IJWMT Vol. 16, No. 5, 8 Oct. 2026
Cover page and Table of Contents: PDF (size: 1313KB)
PDF (1313KB), PP.1-19
Views: 0 Downloads: 0
Boltzmann Exploration, Multicontroller, Load balancing, Reinforcement learning, Software Defined Networking
Software Defined Networking enables centralized control and dynamic programmability, but achieving efficient load balancing across distributed controllers remains challenging because of traffic variability and scalability constraints. Despite the significant potential of reinforcement learning for multicontroller load balancing, existing studies primarily focus on developing new learning architectures or switch migration mechanisms, with limited attention given to systematically comparing exploration strategies under identical network conditions. The experimental setup consists of twelve switches for data forwarding and three distributed controllers following east-west communication. This study compares five exploration strategies of Q-learning using throughput, latency, controller response time, load balancing efficiency, fairness, scalability, and congestion-related metrics. Existing studies also often neglect fault tolerance, energy efficiency, and real-world validation. Under congested conditions, SOFT Q-learning records the lowest throughput reduction of 0.92, followed by upper confidence bound with 1.56, epsilon-greedy strategy with 2.34, prioritized experience replay with 3.29, and Boltzmann exploration with 6.06. The epsilon-greedy strategy achieves the highest throughput of 7.88 and fairness of 0.9999, soft reinforcement learning records the lowest latency of 0.0115 seconds, upper confidence bound achieves the fastest controller response time of 0.0285 seconds, and Boltzmann exploration attains the highest load balance ratio of 0.9865.
Binod Sapkota, Utkarsha Shukla, Babu R. Dawadi, Shashidhar R. Joshi, "Comparative Study of Exploration Strategies in Q- Learning for Multi-Controller SDN Load Balancing", International Journal of Wireless and Microwave Technologies(IJWMT), Vol.16, No.5, pp. 1-19, 2026. DOI:10.5815/ijwmt.2026.05.01
[1]A. H. Abdi., L. Audah., A. Salh., M. A. Alhartomi., H. Rasheed., S. Ahmed., & A. Tahir. (2024). Security control and data planes of SDN: A comprehensive review of traditional, AI, and MTD approaches to security solutions. IEEE Access. DOI: https://doi.org/10.1109/ACCESS.2024.3393548
[2]O. Tkachova., U. Chinaobi., & A. R. Yahya. (2016). A load balancing algorithm for SDN. Scholars Journal of Engineering and Technology (SJET). DOI: https://doi.org/10.21276/sjet.2016.4.11.2
[3]G. Beissenova., A. Zhidebayeva., Z. Kopzhassarova., P. Kozhabekova., B. Myrzakhmetova., M. Kerimbekov., D. Ussipbekova., & N. Yeshenkozhaev. (2024). Load balancing in DCN servers through software defined network machine learning. International Journal of Advanced Computer Science and Applications (IJACSA). DOI: https://doi.org/10.14569/IJACSA.2024.0150254
[4]S. Raghul., T. Subashri., & K. R. Vimal. (2017). Literature survey on traffic-based server load balancing using SDN and OpenFlow. 2017 4th International Conference on Signal Processing, Communications and Networking (ICSCN). DOI: https://doi.org/10.1109/ICSCN.2017.8085416
[5]M. Xiang., M. Chen., D. Wang., & Z. Luo. (2022). Deep reinforcement learning-based load balancing strategy for multiple controllers in SDN. e-Prime – Advances in Electrical Engineering, Electronics and Energy. DOI: https://doi.org/10.1016/j.prime.2022.100038
[6]J. Ali., R. H. Jhaveri., M. Alswailim., & B.-H. Roh. (2023). Escalb: An effective slave controller allocation-based load balancing scheme for multi-domain SDN-enabled-IoT networks. Journal of King Saud University – Computer and Information Sciences. DOI: https://doi.org/10.1016/j.jksuci.2023.101566
[7]W. Negera., F. Schwenker., T. Debelee., H. Melaku., & Y. Ayano. (2022). Review of botnet attack detection in SDN-enabled IoT using machine learning. Sensors. DOI: https://doi.org/10.3390/s22249837
[8]M. Zolotukhin., S. Kumar., & T. Hämäläinen. (2020). Reinforcement learning for attack mitigation in SDN-enabled networks. 2020 6th IEEE Conference on Network Softwarization (NetSoft). DOI: https://doi.org/10.1109/NetSoft48620.2020.9165383
[9]P. Sun., Z. Guo., G. Wang., J. Lan., & Y. Hu. (2020). Marvel: Enabling controller load balancing in software-defined networks with multi-agent reinforcement learning. Computer Networks. DOI: https://doi.org/10.1016/j.comnet.2020.107230
[10]V. Srivastava., & R. S. Pandey. (2021). Machine intelligence approach: To solve load balancing problem with high quality of service performance for multi-controller based software-defined network. Sustainable Computing: Informatics and Systems. DOI: https://doi.org/10.1016/j.suscom.2021.100511
[11]V. Tosounidis., G. Pavlidis., & I. Sakellariou. (2020). Deep Q-learning for load balancing traffic in SDN networks. 11th Hellenic Conference on Artificial Intelligence. DOI: https://doi.org/10.1145/3411408.3411423
[12]D. Tennakoon., S. Karunarathna., & B. Udugama. (2018). Q-learning approach for load-balancing in software-defined networks. 2018 Moratuwa Engineering Research Conference (MERCon).
[13]J. Chen, Y. Wang, J. Ou, C. Fan, X. Lu, C. Liao, X. Huang, and H. Zhang, “Albrl: Automatic load-balancing architecture based on reinforcement learning in software-defined networking,” Wireless Communications and Mobile Computing, vol. 2022, pp. 1–17, 2022, open Access
[14]S. Liang., W. Jiang., F. Zhao., & F. Zhao. (2020). Load balancing algorithm of controller based on SDN architecture under machine learning. Journal of Systems Science and Information. DOI: https://doi.org/10.21078/JSSI-2020-578-11
[15]U. Ahmed., J. C.-W. Lin., & G. Srivastava. (2022). A resource allocation deep active learning based on load balancer for network intrusion detection in SDN sensors. Computer Communications. DOI: https://doi.org/10.1016/j.comcom.2021.12.009
[16]S. Yeo., Y. Naing., T. Kim., & S. Oh. (2021). Achieving balanced load distribution with reinforcement learning-based switch migration in distributed SDN controllers. Electronics. DOI: https://doi.org/10.3390/electronics10020162
[17]M. D. Tache., O. Pascutoiu., & E. Borcoci. (2024). Optimization algorithms in SDN: Routing, load balancing, and delay optimization. Applied Sciences. DOI: https://doi.org/10.3390/app14145967
[18]T. Semong., T. Maupong., S. Anokye., K. Kehulakae., S. Dimakatso., G. Boipelo., & S. Sarefo. (2020). Intelligent load balancing techniques in software-defined networks: A survey. Electronics. DOI: https://doi.org/10.3390/electronics9071091
[19]A. Kumar., & D. Anand. (2021). Load balancing for software-defined network using machine learning. Turkish Journal of Computer and Mathematics Education.
[20]H. Wang., H. Xu., L. Huang., J. Wang., & X. Yang. (2018). Load-balancing routing in software-defined networks with multiple controllers. Computer Networks. DOI: https://doi.org/10.1016/j.comnet.2018.05.012
[21]J. Singh., P. Singh., E. M. Amhoud., & M. Hedabou. (2022). Energy-efficient and secure load balancing technique for SDN-enabled fog computing. Sustainability. DOI: https://doi.org/10.3390/su141912951
[22]S. M. Kerner. What is q-learning? [Online]. Available: https://www.techtarget.com/searchenterpriseai/definition/Qlearning
[23]Q-learning guide: Begin with reinforcement learning basics. Simplilearn. [Online]. Available: https://www.simplilearn.com/tutorials/machine-learning-tutorial/what-is-q-learning
[24]A. Saxena. (2024). Q-learning in machine learning: Explained by experts. Applied AI Course. Available: https://www.appliedaicourse.com/blog/q-learning-in-machine-learning/
[25]R. S. Sutton., & A. G. Barto. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.
[26]D. Takeshi. (2019). Understanding prioritized experience replay. Seita’s Place. Available: https://danieltakeshi.github.io/2019/07/14/per/
[27]T. Schaul., J. Quan., I. Antonoglou., & D. Silver. (2016). Prioritized experience replay. arXiv. Available: https://arxiv.org/abs/1511.05952v4
[28]J.-S. Boon., & X. Zhang. (2022). Lower PAC bound on upper confidence bound-based Q-learning with examples. University of Wisconsin–Madison. Available: https://pages.cs.wisc.edu/~xiaominz/projs/cs761.pdf
[29]T. Haarnoja., H. Tang., P. Abbeel., & S. Levine. (2017). Reinforcement learning with deep energy-based policies. arXiv. Available: https://arxiv.org/abs/1702.08165
[30]M. Goodwin. (2023). What is latency? IBM. Available: https://www.ibm.com/think/topics/latency
[31]A. M. Abdelmoniem., B. Bensaou., & A. J. Abu. (2017). SICC: SDN-based incast congestion control for data centers. Proceedings of the IEEE International Conference on Communications (ICC). DOI: https://doi.org/10.1109/ICC.2017.7996826