Work place: Department of Electronics and Computer Engineering, Pulchowk Campus, IoE, TU, Kathmandu, Nepal
E-mail: srjoshi@ioe.edu.np
Website: https://orcid.org/0000-0001-9480-0495
Research Interests:
Biography
Prof. Dr. Shashidhar R. Joshi received BE in Electrical Engineering degree (Hons.) from Sardar Vallabhbhai Regional College of Engineering, South Gujarat University, Surat, India, in 1984, MSc in electrical engineering from the University of Calgary, Canada, in 1992, and the PhD degree from the Institute of engineering (IoE), Tribhuvan University (TU), Nepal, in 2007. He was a Research Fellow with Osaka Sangyo University, Japan, from 1998 to 1999. He joined the Institute of Engineering in 1985. He is currently a Full Professor in the Department of Electronics and Computer Engineering, IoE, Pulchowk Campus, and was a Visiting Professor with South Asian University, New Delhi, India. He is the former Dean of the IoE, TU. His research interests include image processing, networking, and artificial intelligence.
By Binod Sapkota Utkarsha Shukla Babu R. Dawadi Shashidhar R. Joshi
DOI: https://doi.org/10.5815/ijwmt.2026.05.01, Pub. Date: 8 Oct. 2026
Software Defined Networking enables centralized control and dynamic programmability, but achieving efficient load balancing across distributed controllers remains challenging because of traffic variability and scalability constraints. Despite the significant potential of reinforcement learning for multicontroller load balancing, existing studies primarily focus on developing new learning architectures or switch migration mechanisms, with limited attention given to systematically comparing exploration strategies under identical network conditions. The experimental setup consists of twelve switches for data forwarding and three distributed controllers following east-west communication. This study compares five exploration strategies of Q-learning using throughput, latency, controller response time, load balancing efficiency, fairness, scalability, and congestion-related metrics. Existing studies also often neglect fault tolerance, energy efficiency, and real-world validation. Under congested conditions, SOFT Q-learning records the lowest throughput reduction of 0.92, followed by upper confidence bound with 1.56, epsilon-greedy strategy with 2.34, prioritized experience replay with 3.29, and Boltzmann exploration with 6.06. The epsilon-greedy strategy achieves the highest throughput of 7.88 and fairness of 0.9999, soft reinforcement learning records the lowest latency of 0.0115 seconds, upper confidence bound achieves the fastest controller response time of 0.0285 seconds, and Boltzmann exploration attains the highest load balance ratio of 0.9865.
[...] Read more.Subscribe to receive issue release notifications and newsletters from MECS Press journals