Work place: Department of Electronics and Computer Engineering, Thapathali Campus, IoE, TU, Kathmandu, Nepal
E-mail: replybinod@gmail.com
Website: https://orcid.org/0009-0009-9418-0092
Research Interests:
Biography
Dr. Binod Sapkota earned BE in Electronics and Communication Engineering (2009), MSc in Information System Engineering (2013), and PhD in Computer Engineering (2026) from the Institute of engineering (IoE), Tribhuvan University (TU), Nepal. Since 2014, he has served as an Assistant Professor in the Department of Electronics and Computer Engineering, Thapathali Campus, IoE, TU. He teaches electronics and computer engineering courses and supervises bachelor’s and graduate projects and theses. His research interests include next-generation networking, software-defined networking, AI/ML, cybersecurity, and digital forensics. He has many research papers in the international indexed journals and published books of computer Science and Engineering.
By Binod Sapkota Utkarsha Shukla Babu R. Dawadi Shashidhar R. Joshi
DOI: https://doi.org/10.5815/ijwmt.2026.05.01, Pub. Date: 8 Oct. 2026
Software Defined Networking enables centralized control and dynamic programmability, but achieving efficient load balancing across distributed controllers remains challenging because of traffic variability and scalability constraints. Despite the significant potential of reinforcement learning for multicontroller load balancing, existing studies primarily focus on developing new learning architectures or switch migration mechanisms, with limited attention given to systematically comparing exploration strategies under identical network conditions. The experimental setup consists of twelve switches for data forwarding and three distributed controllers following east-west communication. This study compares five exploration strategies of Q-learning using throughput, latency, controller response time, load balancing efficiency, fairness, scalability, and congestion-related metrics. Existing studies also often neglect fault tolerance, energy efficiency, and real-world validation. Under congested conditions, SOFT Q-learning records the lowest throughput reduction of 0.92, followed by upper confidence bound with 1.56, epsilon-greedy strategy with 2.34, prioritized experience replay with 3.29, and Boltzmann exploration with 6.06. The epsilon-greedy strategy achieves the highest throughput of 7.88 and fairness of 0.9999, soft reinforcement learning records the lowest latency of 0.0115 seconds, upper confidence bound achieves the fastest controller response time of 0.0285 seconds, and Boltzmann exploration attains the highest load balance ratio of 0.9865.
[...] Read more.Subscribe to receive issue release notifications and newsletters from MECS Press journals