
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net
p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net
p-ISSN: 2395-0072
G. Santhi 1 , A. Janani2 , S. Selvaganapathi 3, K. Raja Hariharan4
1Professor, Dept. of Information Technology, Puducherry Technological University, Puducherry, India 234Under Graduate Students, Dept. of Information Technology, Puducherry Technological University, Puducherry, India
Abstract - Cloud computing environments support a wide range of applications requiring reliable and consistent Quality of Service (QoS). Traditional network management approaches are predominantly reactive and fail to handle dynamic traffic conditions, resulting in congestion, inefficient bandwidth utilization, and degraded service performance. This paper proposes a hybrid intelligent framework for dynamic load balancing and bandwidth allocation in Software Defined Networking (SDN)-based cloud environments. The system leverages the centralized control capabilities of SDN through a Ryu OpenFlow 1.3 controller to monitor real-time network conditions and collect detailed traffic statistics. A stacked Long Short-Term Memory (LSTM) model analyzes temporal traffic patterns and predicts congestion levels with 99.22% zone classification accuracy and a ROC-AUC of 1.0000. Based on these predictions, a Dueling Double Deep Q-Network (DQN) agent selects optimal traffic management actions, including flow rerouting, throttling, traffic prioritization, and queue reconfiguration. Enforcement is achievedthrough OpenFlow flow rule installation, metering, DSCP-based marking, and Hierarchical Token Bucket (HTB) queue scheduling. The integrated closed-loop system reduced peak link utilization by 18.4 percentage points, packet drop events by 87.3%, average latency by 35.4%, and jitter by 39.0% compared to an unmanaged baseline. These results confirm that the integration of predictive modeling with reinforcement learning-based control produces measurable and scalable QoSimprovementsinSDN-basedclouddatacenters.
Key Words: Software Defined Networking, Deep QNetwork, Long Short-Term Memory, Quality of Service, Load Balancing, Bandwidth Allocation, Congestion Prediction, OpenFlow, Cloud Computing, Reinforcement Learning
Cloud computing environments host large-scale, dynamic applications requiring consistent Quality of Service (QoS) in terms of latency, throughput, packet loss, and link utilization. Traditional network management techniques
are predominantly reactive and fail to handle sudden traffic variations effectively, leading to congestion and performancedegradation[1].
Software Defined Networking (SDN) addresses this limitation by decoupling the control plane from the data plane, enabling centralized and programmable network control.WhileSDNimprovesflexibilityandglobalnetwork visibility, it lacks the intelligence required to predict and preventcongestion inhighlydynamiccloud environments [9].
The proposed system integrates SDN with a hybrid machine learning framework. A Ryu OpenFlow 1.3 controller continuously monitors network conditions and collectsreal-timemetricsfromallswitches.AstackedLong Short-Term Memory (LSTM) model analyzes temporal patterns in the collected data to predict impending congestion in advance. Based on these predictions, a Dueling Double Deep Q-Network (DQN) agent selects optimal traffic management actions including rerouting, throttling,andprioritization.Thesedecisionsareenforced using OpenFlow flow rules, metering, and Hierarchical Token Bucket (HTB) queuing mechanisms, forming a closed-loopsystemforproactiveQoSmanagement[6][11].
Cloud computing delivers computing resources including processing power, storage, networking, and software as on-demand services over the internet. Cloud environments are characterized by multi-tenancy, where thousands of applications share the same physical infrastructure simultaneously. This shared nature introduces fundamental challenges: network resources such as bandwidth, switch capacity, and link throughput must be dynamically allocated to meet the varying demands of workloads ranging from latency-sensitive videoconferencingtobulkdatatransfers[14].

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net
SoftwareDefinedNetworkingisanarchitecturalapproach that decouples the network control plane from the data plane. In traditional networks, each switch independently makesforwardingdecisionsusingembeddedcontrollogic, making the network difficult to manage at scale. SDN centralizes this intelligence in a software-based controller that maintains a global view of the entire network topology [11]. Thecontrollercommunicates with network devices through the OpenFlow protocol, installing forwarding rules into switch flow tables and exposing the network state to management applications through northboundinterfaces.
Existing SDN-based traffic management solutions handle individualaspectsofnetworkmanagement monitoring, prediction, or enforcement but do not integrate all capabilities within a single closed-loop framework. Machine learning models employed for prediction are not connected to live SDN controllers for automated real-time enforcement, and reinforcement learning-based load balancers operate on instantaneous observed state withoutpredictiveinput,limitingtheirabilitytoactbefore congestion occurs [2][6]. The proposed framework addresses this gap by unifying temporal prediction and reinforcement learning-based control within a continuous two-secondfeedbackloop.
A comprehensive review of existing works identifies ten keystudiesacrossfourfunctionalareas:trafficmonitoring and data collection, congestion prediction, intelligent load balancing,andSDNQoSenforcement.
KhudhairandAthab[1]comparedMininet,Ryu,andiperf3 for SDN traffic generation and data collection, identifying iperf3 as most effective for congestion control datasets. However, no machine learning integration or QoS enforcement mechanism was proposed. Wassie et al. [2] constructedanSDNdatasetusingMininetandRyuwith23 flow-level attributes, achieving up to 100% validation accuracy for elephant flow prediction using XGBoost; however, the prediction model was not integrated into a live controller for real-time enforcement. Han et al. [3] employed Mininet and Ryu for real-time DDoS detection using CICIDS datasets but focused exclusively on attack detectionwithoutQoSoptimization.
Somsuk et al. [4] proposed predictive bandwidth control using clustered LSTM in SDN environments, enabling proactive allocation adjustments under varying traffic loads, but did not incorporate reinforcement learning or
p-ISSN: 2395-0072
multi-pathloadbalancing.Agrawaletal.[5]demonstrated AI-based network fault and QoS degradation prediction using LSTM, but without an SDN enforcement layer or bandwidth reallocation mechanism. Khalid et al. [6] appliedQ-learningandDQNforadaptiveloadbalancingin high-load SDN scenarios but lacked dynamic bandwidth allocationandpredictiveinputpriortocongestiononset.
Kirti et al. [7] reviewed fault tolerance techniques including replication and checkpointing for distributed environments, but these are predominantly reactive with no SDN-integrated proactive control. Tamilarasu et al. [8] optimized cloud QoS through multi-objective Particle Swarm Optimization but without network-level SDN enforcement or dynamic load balancing. Thazin and Nwe [9]reviewedQoScapabilitiesofOpenFlowincludingmeter tables, HTB queuing, and DiffServ; however, QoS rules are configured statically without real-time congestion intelligence. Imran et al. [10] detected unauthorized DSCP modifications using deep learning models with 99.28% accuracy, but the study focused on detection without automaticcorrectiveenforcement.
Ref. Techniqu e Key Contribution Research Gap
[1] Mininet, Ryu, iperf3 SDNdata collection benchmark NoMLintegration orQoS enforcement
[2] XGBoost, GBM Elephantflow prediction(100% accuracy) Notintegrated withliveSDN controller
[3] MLon CICIDS Real-timeDDoS detectioninSDN Nocongestion predictionorload balancing
[4] Clustered LSTM Predictive bandwidthcontrol inSDN NoRLor congestionzone classification
[5] LSTM,ML models AI-basedfaultand QoSprediction NoSDN enforcementor reallocationlayer
[6] Qlearning, DQN Adaptiveload balancingunder high Notraffic predictionor bandwidth allocation
[7] Review: replication , checkpoin Faulttolerancefor distributedcloud Reactiveonly,no proactiveSDN control

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net
ting
[8] Multi-obj. PSO CloudQoSvia resource scheduling
[9] OpenFlow QoS review
HTB,metering, DiffServinSDN
[10] CNN,RNN, LSTM DSCP manipulation detection
NoSDNnetworklevelenforcement
Staticrules,no real-time intelligence
Detectiononly,no corrective enforcement
The synthesis of these studies reveals that no existing work integrates real-time traffic monitoring, time-series congestion prediction, reinforcement learning-based decision making, and multi-mechanism QoS enforcement within a single closed-loop SDN framework. This gap constitutes the core motivation and contribution of the proposedsystem.
The proposed system follows a closed-loop architecture comprising four tightly coupled modules: Network Traffic Monitoring and Data Collection, Traffic Analysis and Congestion Prediction, Intelligent Load Balancing and Bandwidth Optimization, and SDN Control and QoS Enforcement. The system operates on a two-second feedback cycle, continuously adapting to dynamic traffic conditionswithoutmanualintervention.
3.1 System Architecture
At the network layer, a Mininet-emulated SDN tree topologywithfanout=2anddepth=5produces31switches and 32 hosts, with all links configured at 100 Mbps and 2 msdelayusingTCLink.Diversetraffictypes Web,VoIP, Gaming, IoT, and Cloud API flows are generated using iperf3 based on a 1,000-flow simulation profile derived fromtheCICIDS2017dataset[12].
The Ryu OpenFlow 1.3 controller serves as the central intelligencehub,pollingall31switcheseverytwoseconds and collecting 30 metrics per port including throughput, utilization, latency, jitter, and packet drop deltas. Custom LLDP probe frames measure one-way delay per interswitch link, with jitter maintained as an exponentially weightedmovingaveragewithsmoothingfactorα=0.2.
p-ISSN: 2395-0072

3.2
The Canadian Institute for Cybersecurity CICIDS 2017 datasetservesasthebasisforthetrafficsimulationprofile. The dataset contains 3,594,396 flow records across 79 features per flow and occupies 968 MB. Flow-level attributes are used to derive realistic traffic patterns across five application classes, enabling representative simulationofproductioncloudtrafficconditions.
Module I establishes the foundation of the proposed system by creating the SDN simulation environment, generating realistic network traffic, and collecting realtime network metrics. The Ryu controller employs a staggered polling strategy to prevent controller overload each switch is queried at evenly spaced sub-intervals withinthetwo-secondpollingwindow.Acongestiondualsignaldetectionmechanismclassifieseachportintooneof four congestion zones based on the combination of utilizationandpacketdropsignals.
Condition
Assigned Label
Bothutilization(>95%)and dropsignalsactive Critical
Eithersignalaloneactive Congested
Utilization70–95%,nodrops Warning Allotherconditions Normal
The final dataset contains 74,336 rows across 92 unique switch ports with 30 columns and zero missing values. Congestion was distributed as 62.0% normal, 5.6%

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net
warning,and32.4%congestedportobservationsacrossall five traffic classes. Of the 1,000 scheduled flows, 910 completed successfully (91.0% flow success rate), and all REST API endpoints serving telemetry to downstream modules operated without failure throughout the simulation.
Module II analyzes the network metrics collected in Module I and predicts congestion using a deep learning approach.Theraw30-columnCSVlogispreprocessedtoa 17-featureinputmatrixbyremovingidentitycolumns,raw byte counters, zero-variance columns, and redundant derived features. Bandwidth values are clipped at 500 Mbps to eliminate counter wrap spikes, and a StandardScaler normalizes each feature to zero mean and unitvariance.
The 17-feature time series for each of the 92 monitored switch ports is segmented into sliding windows of length 10 with a stride of 1, capturing 20 seconds of consecutive port observations per window. Windows are constructed perportinstrictchronologicalordertopreservetemporal structure and labeled with the zone and congestion probability of the final timestep, producing 73,508 windowswithnomissingorinfinitevalues.
The model is a two-layer stacked LSTM with 128 hidden units per layer, designed to capture temporal dependencies in port-level traffic sequences. Batch Normalization is applied after the final LSTM layer before theoutputispassedtotwoindependentpredictionheads. HeadAperformsthree-classcongestionzoneclassification (normal, warning, congested) using a Linear(128→64)→ReLU→Dropout→Linear(64→3) architecture with Weighted CrossEntropy loss. Head B outputs a scalar congestion probability using a Linear(128→32)→ReLU→Dropout→Linear(32→1) architecture with BCEWithLogits loss and positive class weight2.09toaddressclassimbalance.Thetotaltrainable parametercountis220,228.
Table 3: LSTM Model Architecture Specification
Component Specification
LSTMLayers 2(Stacked),128hiddenunitsper layer
2395-0072
DropoutRate 0.3(betweenlayers)
Post-LSTM Normalization BatchNormalization
HeadA Zone Classification
HeadB Congestion Probability
Linear(128→64)→ReLU→Dropout→ Linear(64→3)
Linear(128→32)→ReLU→Dropout→ Linear(32→1)
TotalTrainable Parameters 220,228
Optimizer
Adam(lr=0.001)
BestEpoch 26,ValidationLoss=0.008175
The dataset is split at the port level rather than the window level to prevent temporal data leakage, with 64 ports (51,136 windows) used for training, 14 ports (11,186 windows) for validation, and 14 ports (11,186 windows) for final evaluation. The model was trained for 36epochsonCPUin333seconds.
Table 4: Zone Classification Performance (Head A)

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
Table 5: Binary Congestion Detection Performance (Head B)

FalsePositives 0
The model achieved zero false positives for binary congestion detection. A ROC-AUC of 1.0000 confirms that the probability output of Head B is well-calibrated for downstreamusebytheDQNagent.Allclassificationerrors occurred at the normal/warning and congested/warning boundaries, confirming a sharp decision boundary betweenhealthyandcongestednetworkstates.

Classification Accuracy
Module III implements intelligent decision-making for dynamic traffic management using deep reinforcement learning. A Dueling Double Deep Q-Network agent observes a 12-feature state vector per switch aggregated from LSTM state vectors of all ports belonging tothatswitch andselectsoneoffivediscreteactionsto maintainQoSacrosstheSDNtopology.
State features capture congestion severity (maximum and mean P(congested)), bandwidth availability (minimum headroom), latency, packet loss, jitter, aggregate packet drop increments, rolling utilization trends, neighbouring switch utilization, and the fraction of congested ports. All features are normalized to [0, 1] using fixed maximum values. The five available actions are: Do Nothing (no configuration change), Reroute (OpenFlow Flow Mod at priority200redirectingtotheleast-loadedalternateport), Throttle(OpenFlowMeterimposinga5Mbpscaponbesteffort traffic), Prioritise (DSCP value 46 Expedited Forwarding applied via set-field action), and Reset (cookie-basedmassremovalofallagent-installedrulesvia cookie0xDEADBEEF).

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net
The DQN uses a Dueling architecture with a shared trunk of two fully connected layers (128 units, ReLU activation) that branches into a Value stream V(s) and an Advantage stream A(s, a). The final Q-value is computed as Q(s, a) = V(s) + A(s, a) − mean(A(s, :)). This decomposition allows theagenttoindependentlylearnthegeneraldesirabilityof a state and the relative advantage of each action, which is particularly effective in SDN environments where many actions have similar expected returns on calm switches. Double DQN further reduces Q-value overestimation by usingthemainnetworkforactionselectionandthetarget networkforactionevaluation.
The reward signal is computed from real network measurements fetched after each enforcement action: R = 1.5 × throughput_norm − 1.0 × latency_norm − 2.0 × loss_norm − 0.5 × jitter_norm + 1.0 × congestion_reduction_bonus − 0.3 × idle_penalty. Packet loss carries the highest penalty weight given its critical impact across all traffic classes. A bonus is awarded when the fraction of congested ports decreases between consecutive decision steps, reinforcing effective interventions. An idle penalty discourages unnecessary actionsonnon-congestedswitches.
TotalTrainingSteps 1,000
ReplayMemoryCapacity 10,000transitions
BatchSize
EpsilonStart/End 1.0/0.05(decay:0.997)
TargetNetworkSync Every20steps
GradientClipping max_norm=10.0
Trainingprogressedthroughthreedistinctphases.During exploration (steps 0–200), random actions built a diverse experiencebuffer.Asepsilondecayedbelow0.5fromstep 200 onward, the agent began preferring do_nothing on calm switches and selecting reroute or throttle when P(congested) exceeded 0.7. By step 500, the policy had
2395-0072
stabilised with a consistent upward trend in cumulative rewardthroughtheexploitationphase.


6: DQN Agent Cumulative Reward Progression Across 1,000 Training Steps
Module IV enforces the traffic management decisions generated by the DQN agent directly in the SDN environment. The Ryu controller translates each selected action into concrete network modifications using four complementaryOpenFlowmechanisms.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net
OpenFlow 1.3 flow rules for rerouting are installed at priority 200, overriding default L2 learning rules at priority1,andredirecttraffictotheleast-loadedalternate output port. All agent-installed rules use idle_timeout=0 and hard_timeout=0, making them permanent until explicitlyremovedbytheresetaction,whichissuesaflow delete message matching cookie=0xDEADBEEF across all affected switches. OpenFlow meter tables with a single drop band enforce the 5 Mbps rate cap on best-effort traffic under the throttle action, applying only to flows identified by theirDSCPvalueto ensurethat high-priority flows bypass metering entirely. DSCP value 46 (Expedited Forwarding per-hop behaviour) is applied through the prioritise action, marking packets for minimum-delay and minimum-droptreatmentateverydownstreamhop.
Three HTB queues are configured per Open vSwitch port via the OVSDB interface: Queue 0 (high priority, VoIP and Gaming, 60 Mbps guaranteed), Queue 1 (medium priority, Web and Cloud API, 30 Mbps guaranteed), and Queue 2 (low priority, IoT and Bulk Transfer, 10 Mbps). The DQN agent's throttle and prioritise actions trigger corresponding changes to meter and queue assignments, ensuring differentiated service delivery across all traffic classes.
TheLSTM-onlyconditionisidenticaltothebaselineacross all metrics, confirming that the enforcement layer is essential prediction alone has no effect on network behavior. The full proposed system reduced peak link utilization by 18.4 percentage points, packet drop events by 87.3%, average latency by 35.4%, and jitter by 39.0%. With HTB enforcement, VoIP/Gaming throughput improved from 27.7 Mbps under FIFO to 59.7 Mbps a 115% increase for the highest-priority traffic class while IoT and bulk traffic was constrained to 10 Mbps under the throttle action, eliminating drop overhead and recoveringlinkheadroomfrom12.7%to28.8%.
Table 9: Feature-Level Comparison with Existing Approaches
columndataset
(zone+ probability)
End-to-end system performance is evaluated across three experimental conditions: the unmanaged baseline Ryu controller, LSTM prediction active with no enforcement, andthefullproposedsystemwithallfourmodulesactive.
8: End-to-End QoS Comparison Across Experimental Conditions
(hashbased)
+DSCP+ Metering

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072



This paper presented a hybrid intelligent framework for dynamic load balancing and bandwidth allocation in SDNbased cloud environments. The system integrates four tightly coupled modules real-time network monitoring, LSTM-based congestion prediction, DQN-based decisionmaking, and OpenFlow enforcement into a continuous closed-loop that adapts proactively to changing traffic conditionswithoutmanualintervention.
The LSTM model achieved 99.22% zone classification accuracy and a ROC-AUC of 1.0000 for binary congestion detection with zero false positives. The Dueling Double DQN agent converged to a stable traffic management policy within 1,000 training steps, resolving 87.3% of congestion episodes. End-to-end evaluation demonstrated reductions of 18.4 percentage points in peak link utilization,87.3%inpacketdropevents,35.4%inaverage latency, and 39.0% in jitter compared to the unmanaged baseline. These results confirm that the integration of temporal prediction with reinforcement learning-based control produces qualitatively different and measurably superior QoS outcomes compared to reactive or partially integratedalternatives.
Future work will explore multi-agent reinforcement learning for distributed network segment management, integration with production-level SDN controllers such as ONOS or OpenDaylight, advanced prediction architectures including Transformers and hybrid deep learning models, security-aware traffic management incorporating intrusion detection, and energy-efficient resource allocation strategies aligned with sustainable cloud operationobjectives.
The authors acknowledge the open-source communities behind Mininet, Open vSwitch, the Ryu SDN Framework, and PyTorch for providing the tools that enabled the simulation and implementation of the proposed system. The Canadian Institute for Cybersecurity is acknowledged for making the CICIDS 2017 dataset publicly available for researchpurposes.
[1]T.KhudhairandO.A.Athab,"Recenttoolsofsoftwaredefined networking traffic generation and collection," Al-Khwarizmi Engineering Journal, vol. 21, no. 2, pp. 93–105,2025.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
[2]G.Wassie,J.Ding,andY.Wondie,"Trafficpredictionin SDN for explainable QoS using deep learning approach," Scientific Reports, vol. 13, Art. no. 20607, Nov.2023,doi:10.1038/s41598-023-47999-3.
[3] D. Han, H. Li, X. Fu, and S. Zhou, "Traffic feature selection and distributed denial of service attack detection in software-defined networks based on machine learning," Sensors, vol. 24, no. 13, Art. no. 4344,2024,doi:10.3390/s24134344.
[4] K. Somsuk, S. Khummanee, and P. Songram, "Dynamic predictive feedback mechanism for intelligent bandwidthcontrol in future SDNnetworks," Network, vol. 5, no. 1, Art. no. 3, 2025, doi: 10.3390/network5010003.
[5] P. Agrawal, K. Chhillar, S. Shrivastava, and D. Tomar, "AI-powered predictive models for network fault detection and proactive QoS management: A comprehensive analysis," International Journal of SciencesandInnovationEngineering,vol.2,no.10,pp. 227–244,2025.
[6] M. Khalid, H. M. Muhi-Aldeen, and B. R. M. Alhamdani, "New learning approach for high-load traffic optimizationinsoftware-definednetworks,"Journalof Intelligent SystemsandInternetof Things,vol.17,no. 1,pp.255–270,2025.
[7]M.Kirti,A.K.Maurya,andR.S.Yadav,"Fault-tolerance approaches for distributed and cloud computing environments: A systematic review, taxonomy and future directions," Concurrency and Computation: PracticeandExperience,vol.36,no.13,Art.no.e8081, 2024,doi:10.1002/cpe.8081.
[8] P. Tamilarasu, G. Singaravel, P. Manoharan, and S. Selvarajan, "QoS transformation in the cloud: Advancingservicequalitythroughinnovativeresource scheduling," IET Communications, vol. 19, no. 1, Art. no.e70020,2025,doi:10.1049/cmu2.70020.
[9]N.ThazinandK.M.Nwe,"Qualityofserviceinsoftware defined network: Leveraging OpenFlow protocol," Journal of Information Systems Engineering and Management, vol. 9, no. 2, Art. no. 24127, 2024, doi: 10.55267/jisem.v9i2.24127.
[10] A. Imran et al., "Detection of DSCP-based traffic prioritization manipulations and their impact on network performance," Scientific Reports, vol. 15, Art. no.3080,2025,doi:10.1038/s41598-025-87463-w.
[11]N.McKeownetal.,"OpenFlow:Enablinginnovationin campus networks," ACM SIGCOMM Computer
Communication Review, vol. 38, no. 2, pp. 69–74, Apr. 2008,doi:10.1145/1355734.1355746.
[12] B. Lantz, B. Heller, and N. McKeown, "A network in a laptop: Rapid prototyping for software-defined networks," in Proc. 9th ACM SIGCOMM Workshop on Hot Topics in Networks (HotNets-IX), Monterey, CA, USA, Oct. 2010, pp. 1–6, doi: 10.1145/1868447.1868466.
[13]N.Gudeetal.,"NOX:Towardsanoperatingsystemfor networks," ACM SIGCOMM Computer Communication Review, vol. 38, no. 3, pp. 105–110, Jul. 2008, doi: 10.1145/1384609.1384625.
[14] M. Al-Fares, A. Loukissas, and A. Vahdat, "A scalable, commoditydatacenternetworkarchitecture,"inProc. ACM SIGCOMM, Seattle, WA, USA, Aug. 2008, pp. 63–74,doi:10.1145/1402958.1402967.
[15]M.Al-Fares,S.Radhakrishnan,B.Raghavan,N.Huang, and A. Vahdat, "Hedera: Dynamic flow scheduling for data center networks," in Proc. 7th USENIX NSDI, San Jose,CA,USA,Apr.2010,pp.89–92.
[16] S. Hochreiter and J. Schmidhuber, "Long short-term memory,"NeuralComputation,vol.9,no.8,pp.1735–1780,Nov.1997,doi:10.1162/neco.1997.9.8.1735.
[17] V. Mnih et al., "Human-level control through deep reinforcementlearning,"Nature,vol.518,no.7540,pp. 529–533,Feb.2015,doi:10.1038/nature14236.
[18] H. Van Hasselt, A. Guez, and D. Silver, "Deep reinforcement learning with double Q-learning," in Proc. 30th AAAI Conference on Artificial Intelligence, Phoenix,AZ,USA,Feb.2016,pp.2094–2100.
[19] Z. Wang et al., "Dueling network architectures for deepreinforcementlearning,"inProc.33rdICML,New York,NY,USA,Jun.2016,pp.1995–2003.
[20] S. Blake et al., "An architecture for differentiated services," IETF RFC 2475, Dec. 1998. [Online]. Available:https://www.rfc-editor.org/rfc/rfc2475.