Work place: Department of Electronics and Computer Engineering, Thapathali Campus, IoE, TU, Kathmandu, Nepal
E-mail: utkar24@gmail.com
Website: https://orcid.org/0009-0000-8832-5747
Research Interests:
Biography
Utkarsha Shukla received his B.Tech in Computer Science and Engineering from KL University (KL Deemed to be University) in 2019. He received his MSc. in Informatics and Intelligent Systems Engineering from Institute of Engineering, Thapathali Campus, Tribhuvan University in 2025. He teaches computer science related courses at bachelors’ level, His research interest areas include Software Defined Networking, Natural Language Processing, Machine Learning, Cyber Security, and Digital Forensics.
By Binod Sapkota Utkarsha Shukla Babu R. Dawadi Shashidhar R. Joshi
DOI: https://doi.org/10.5815/ijwmt.2026.05.01, Pub. Date: 8 Oct. 2026
Software Defined Networking enables centralized control and dynamic programmability, but achieving efficient load balancing across distributed controllers remains challenging because of traffic variability and scalability constraints. Despite the significant potential of reinforcement learning for multicontroller load balancing, existing studies primarily focus on developing new learning architectures or switch migration mechanisms, with limited attention given to systematically comparing exploration strategies under identical network conditions. The experimental setup consists of twelve switches for data forwarding and three distributed controllers following east-west communication. This study compares five exploration strategies of Q-learning using throughput, latency, controller response time, load balancing efficiency, fairness, scalability, and congestion-related metrics. Existing studies also often neglect fault tolerance, energy efficiency, and real-world validation. Under congested conditions, SOFT Q-learning records the lowest throughput reduction of 0.92, followed by upper confidence bound with 1.56, epsilon-greedy strategy with 2.34, prioritized experience replay with 3.29, and Boltzmann exploration with 6.06. The epsilon-greedy strategy achieves the highest throughput of 7.88 and fairness of 0.9999, soft reinforcement learning records the lowest latency of 0.0115 seconds, upper confidence bound achieves the fastest controller response time of 0.0285 seconds, and Boltzmann exploration attains the highest load balance ratio of 0.9865.
[...] Read more.Subscribe to receive issue release notifications and newsletters from MECS Press journals