IJWMT Vol. 16, No. 4, 8 Aug. 2026
Cover page and Table of Contents: PDF (size: 601KB)
PDF (601KB), PP.427-436
Views: 0 Downloads: 0
Cloud Computing, Reinforcement Learning, Resource Allocation, Workload Prediction, LSTM, Deep Q-Network, Prediction Confidence
Cloud schedulers that pair workload prediction with reinforcement learning (RL) rarely check whether a given prediction can actually be trusted, and earlier confidence-gated designs often mix current and future information inconsistently. We fix that inconsistency and build a causally consistent, confidence-gated LSTM-DQN scheduler: an LSTM forecasts next-step workload, a retrospective, error-based confidence score gates how much a Deep Q-Network (DQN) scheduler leans on that forecast, and only information available at decision time is ever used. We implement and pilot-test this architecture in a Python-based discrete-event simulation configured to match a CloudSim-style environment (10 hosts, 30 VMs), benchmarking it against FCFS, Round Robin, standard RL, two ablation variants, and two simplified state-of-the-art comparators across five random seeds. The results show the method works as intended: it trains stably and safely on every seed, holds response time and SLA violations in line with standard RL and simple heuristics, and clearly outperforms a metaheuristic-augmented Q-learning baseline, which suffered severe instability under the same conditions. Code, raw results, and statistical tests are released for independent verification, with scaled-up training identified as the natural next step to test whether larger performance gains emerge.
Preeti Puranik, Sushila Sonare, "Online Confidence-Gated LSTM-DQN for Dynamic Cloud Resource Allocation", International Journal of Wireless and Microwave Technologies(IJWMT), Vol.16, No.4, pp. 427-436, 2026. DOI:10.5815/ijwmt.2026.04.25
[1]C. Li, W. Dong, L. He, P. Zhang, and Y. Li, "A Hierarchical Multi-Agent Reinforcement Learning Approach for Intelligent Decision-Making in Simulation Environment," in 2025 2nd International Symposium on AI and Cybersecurity (ISAICS), IEEE, Oct. 2025, pp. 1–5. doi: 10.1109/ISAICS66888.2025.11349954.
[2]S. Kayalvili, R. Senthilkumar, S. Yasotha, and R. S. Kamalakannan, "An optimized resource allocation in cloud using prediction enabled reinforcement learning," Sci. Rep., vol. 15, no. 1, p. 36088, Oct. 2025, doi: 10.1038/s41598-025-19927-2.
[3]B. Kruekaew and W. Kimpan, "Multi-Objective Task Scheduling Optimization for Load Balancing in Cloud Computing Environment Using Hybrid Artificial Bee Colony Algorithm with Reinforcement Learning," IEEE Access, vol. 10, pp. 17803–17818, 2022, doi: 10.1109/ACCESS.2022.3149955.
[4]S. Tuli, S. Ilager, K. Ramamohanarao, and R. Buyya, "Dynamic Scheduling for Stochastic Edge-Cloud Computing Environments Using A3C Learning and Residual Recurrent Neural Networks," IEEE Trans. Mob. Comput., vol. 21, no. 3, pp. 940–954, Mar. 2022, doi: 10.1109/TMC.2020.3017079.
[5]Y. Chen, Z. Liu, Y. Zhang, Y. Wu, X. Chen, and L. Zhao, "Deep Reinforcement Learning-Based Dynamic Resource Management for Mobile Edge Computing in Industrial Internet of Things," IEEE Trans. Industr. Inform., vol. 17, no. 7, pp. 4925–4934, Jul. 2021, doi: 10.1109/TII.2020.3028963.
[6]L. Schuler, S. Jamil, and N. Kuhl, "AI-based Resource Allocation: Reinforcement Learning for Adaptive Auto-scaling in Serverless Environments," in 2021 IEEE/ACM 21st International Symposium on Cluster, Cloud and Internet Computing (CCGrid), IEEE, May 2021, pp. 804–811. doi: 10.1109/CCGrid51090.2021.00098.
[7]Y. He, Y. Wang, C. Qiu, Q. Lin, J. Li, and Z. Ming, "Blockchain-Based Edge Computing Resource Allocation in IoT: A Deep Reinforcement Learning Approach," IEEE Internet Things J., vol. 8, no. 4, pp. 2226–2237, Feb. 2021, doi: 10.1109/JIOT.2020.3035437.
[8]W. Guo, W. Tian, Y. Ye, L. Xu, and K. Wu, "Cloud Resource Scheduling With Deep Reinforcement Learning and Imitation Learning," IEEE Internet Things J., vol. 8, no. 5, pp. 3576–3586, Mar. 2021, doi: 10.1109/JIOT.2020.3025015.
[9]M. T. Islam, S. Karunasekera, and R. Buyya, "Performance and Cost-Efficient Spark Job Scheduling Based on Deep Reinforcement Learning in Cloud Computing Environments," IEEE Transactions on Parallel and Distributed Systems, vol. 33, no. 7, pp. 1695–1710, Jul. 2022, doi: 10.1109/TPDS.2021.3124670.
[10]A. Belgacem, "Dynamic resource allocation in cloud computing: analysis and taxonomies," Computing, vol. 104, no. 3, pp. 681–710, Mar. 2022, doi: 10.1007/s00607-021-01045-2.
[11]T. Goyal, A. Singh, and A. Agrawal, "Cloudsim: simulator for cloud computing infrastructure and modeling," Procedia Eng., vol. 38, pp. 3566–3572, 2012, doi: 10.1016/j.proeng.2012.06.412.
[12]R. Buyya, R. Ranjan, and R. N. Calheiros, "Modeling and simulation of scalable Cloud computing environments and the CloudSim toolkit: Challenges and opportunities," in 2009 International Conference on High Performance Computing & Simulation, IEEE, Jun. 2009, pp. 1–11. doi: 10.1109/HPCSIM.2009.5192685.
[13]R. N. Calheiros, R. Ranjan, A. Beloglazov, C. A. F. De Rose, and R. Buyya, "CloudSim: a toolkit for modeling and simulation of cloud computing environments and evaluation of resource provisioning algorithms," Softw. Pract. Exp., vol. 41, no. 1, pp. 23–50, Jan. 2011, doi: 10.1002/spe.995.
[14]J. Wilkes, "Google cluster-usage traces v3," Technical Report, Google Inc., Mountain View, CA, USA, 2020. Available: https://github.com/google/cluster-data/blob/master/ClusterData2019.md
[15]H. M. D. Kabir, A. Khosravi, S. K. Mondal, M. Rahman, S. Nahavandi, and R. Buyya, "Uncertainty-Aware Decisions in Cloud Computing: Foundations and Future Directions," ACM Comput. Surv., vol. 54, no. 4, Article 74, pp. 1–30, May 2021, doi: 10.1145/3447583.