Decentralized Meta-Reinforcement Learning with Graphical Neural Networks for Dynamic Spectrum Access in 5G IoT Environments

PDF (756KB), PP.20-36

Views: 0 Downloads: 0

Author(s)

Jayesh Kumar Dabi 1,* Priyadarshi Ashok Dahat 2

1. Department of Electronics & Communication Engineering, SVCE Indore, RGPV, India

2. Department of Electronics &Telecommunication Engineering,IET, DAVV Indore, India

* Corresponding author.

DOI: https://doi.org/10.5815/ijwmt.2026.05.02

Received: 21 May 2026 / Revised: 17 Jun. 2026 / Accepted: 14 Jul. 2026 / Published: 8 Oct. 2026

Index Terms

Decentralized Meta-Reinforcement Learning, Dynamic Spectrum Access, Cognitive Radio, 5G IoT, Spectrum Sharing, Multi-Agent System, Adaptive Resource Management

Abstract

5G IoT networks which use cognitive radio technology need dynamic spectrum access (DSA) to achieve fast response times during extreme environmental changes while managing extensive network operations. The research introduces a decentralized graph-based meta-reinforcement learning framework which enables cognitive IoT devices to learn spectrum access methods through decentralized learning. The proposed method uses Model-Agnostic Meta-Learning (MAML) with Graph Neural Networks (GNNs) to enable structure-aware few-shot adaptation which requires only local observations and neighbor interactions to function. Agents require between 1 and 5 gradient steps to develop new spectrum adaptation capabilities. The proposed framework achieved 83% spectrum utilization, an average throughput of 1.6 packets/slot, Jain's fairness index of 0.92, and a collision rate of 5%, outperforming Meta-RL and GCN-RL baselines. The study demonstrates that decentralized Meta-RL with relational learning offers an effective and scalable method for managing intelligent spectrum in future wireless IoT networks.

Cite This Paper

Jayesh Kumar Dabi, Priyadarshi Ashok Dahat, "Decentralized Meta-Reinforcement Learning with Graphical Neural Networks for Dynamic Spectrum Access in 5G IoT Environments", International Journal of Wireless and Microwave Technologies(IJWMT), Vol.16, No.5, pp. 20-36, 2026. DOI:10.5815/ijwmt.2026.05.02

Reference

[1]Cisco. (2020). Cisco Annual Internet Report (2018–2023) White Paper. Cisco Systems Inc.
[2]Qing Zhao., & Brian M. Sadler. (2007). A survey of dynamic spectrum access: Signal processing, networking, and regulatory policy. IEEE Signal Processing Magazine. DOI: https://doi.org/10.1109/MSP.2007.361604.
[3]Simon Haykin. (2005). Cognitive radio: Brain-empowered wireless communications. IEEE Journal on Selected Areas in Communications. DOI: https://doi.org/10.1109/JSAC.2004.839380.
[4]Ian F. Akyildiz., W.-Y. Lee., M. C. Vuran., & S. Mohanty. (2006). NeXt generation/dynamic spectrum access/cognitive radio wireless networks: A survey. Computer Networks. DOI: https://doi.org/10.1016/j.comnet.2006.05.001.
[5]H. Ye., G. Y. Li., & B.-H. Juang. (2019). Deep reinforcement learning based resource allocation for V2V communications. IEEE Transactions on Vehicular Technology. DOI: https://doi.org/10.1109/TVT.2019.2897134.
[6]W. Bai., G. Zheng., W. Xia., Y. Mu., & Y. Xue. (2025). Multi-user opportunistic spectrum access for cognitive radio networks based on multi-head self-attention and multi-agent deep reinforcement learning. Sensors. DOI: https://doi.org/10.3390/s25072025.
[7]L. Chen., Z. Wang., X. Zhao., X. Shen., & W. He. (2024). A dynamic spectrum access algorithm based on deep reinforcement learning with novel multi-vehicle reward functions in cognitive vehicular networks. Telecommunication Systems. DOI: https://doi.org/10.1007/s11235-024-01188-5.
[8]O. Naparstek., & K. Cohen. (2019). Deep multi-user reinforcement learning for distributed dynamic spectrum access. IEEE Transactions on Wireless Communications. DOI: https://doi.org/10.1109/TWC.2018.2879433.
[9]Y. Zhou., F. Zhou., Y. Wu., R. Q. Hu., & Y. Wang. (2020). Subcarrier assignment schemes based on Q-learning in wideband cognitive radio networks. IEEE Transactions on Vehicular Technology. DOI: https://doi.org/10.1109/TVT.2019.2953809.
[10]J. Foerster., G. Farquhar., T. Afouras., N. Nardelli., & S. Whiteson. (2018). Counterfactual multi-agent policy gradients. Proceedings of the AAAI Conference on Artificial Intelligence. DOI: https://doi.org/10.1609/AAAI.V32I1.11794.
[11]X. Liu, J. Wu, S. Chen, "A context-based meta-reinforcement learning approach to efficient hyperparameter optimization," Neurocomputing, vol. 478, pp. 89-103, Mar. 2022, doi: 10.1016/j.neucom.2021.12.086.
[12]G. Chai., W. Wu., Q. Yang., R. Liu., & F. R. Yu. (2022). Learning-based resource allocation for ultra-reliable V2X networks with partial CSI. IEEE Transactions on Communications. DOI: https://doi.org/10.1109/TCOMM.2022.3199018.
[13]S. Wang., H. Liu., P. H. Gomes., & B. Krishnamachari. (2018). Deep reinforcement learning for dynamic multichannel access in wireless networks. IEEE Transactions on Cognitive Communications and Networking. DOI: https://doi.org/10.1109/TCCN.2018.2809722.
[14]C. Zhong., Z. Lu., M. C. Gursoy., & S. Velipasalar. (2018). Actor-critic deep reinforcement learning for dynamic multichannel access. Proceedings of the IEEE Global Conference on Signal and Information Processing (GlobalSIP). DOI: https://doi.org/10.1109/GlobalSIP.2018.8646405.
[15]J. Huang., J. Wan., B. Lv., Q. Ye., & Y. Chen. (2023). Joint computation offloading and resource allocation for edge-cloud collaboration in internet of vehicles via deep reinforcement learning. IEEE Systems Journal. DOI: https://doi.org/10.1109/JSYST.2023.3249217.
[16]M. Marwani., & G. Kaddoum. (2024). Graph neural networks approach for joint wireless power control and spectrum allocation. IEEE Transactions on Machine Learning in Communications and Networking. DOI: https://doi.org/10.1109/TMLCN.2024.3408723.
[17]H. Song., L. Liu., J. Ashdown., & Y. Yi. (2021). A deep reinforcement learning framework for spectrum management in dynamic spectrum access. IEEE Internet of Things Journal. DOI: https://doi.org/10.1109/JIOT.2021.3052691.
[18]H.-H. Chang., Y. Song., T. T. Doan., & L. Liu. (2023). Federated multi-agent deep reinforcement learning (Fed-MADRL) for dynamic spectrum access. IEEE Transactions on Wireless Communications. DOI: https://doi.org/10.1109/TWC.2022.3233436.
[19]Hainan Qi., Wenjie Ma., & Bingsong Zhao. (2025). Federated heterogeneous multi-agent deep reinforcement learning-based attack resilience scheduling for heterogeneous multi-integrated energy system. Energy Reports. DOI: https://doi.org/10.1016/j.egyr.2025.05.085.
[20]S. Zhang., K.-Y. Lam., B. Shen., L. Wang., & F. Li. (2023). Dynamic spectrum access for internet of things with hierarchical federated deep reinforcement learning. Ad Hoc Networks. DOI: https://doi.org/10.1016/j.adhoc.2023.103257.
[21]X. Tan., L. Zhou., H. Wang., Y. Sun., H. Zhao., B.-C. Seet., J. Wei., & V. C. M. Leung. (2022). Cooperative multi-agent reinforcement learning based distributed dynamic spectrum access in cognitive radio networks. IEEE Internet of Things Journal. DOI: https://doi.org/10.1109/JIOT.2022.3168296.
[22]H. Zhang., N. Liu., X. Chu., K. Long., A.-H. Aghvami., & V. C. M. Leung. (2017). Network slicing based 5G and future mobile networks: Mobility, resource management, and challenges. IEEE Communications Magazine. DOI: https://doi.org/10.1109/MCOM.2017.1600940.
[23]A. Al-Fuqaha., M. Guizani., M. Mohammadi., M. Aledhari., & M. Ayyash. (2015). Internet of things: A survey on enabling technologies, protocols, and applications. IEEE Communications Surveys & Tutorials. DOI: https://doi.org/10.1109/COMST.2015.2444095.
[24]Y. Zeng., J. Xu., & R. Zhang. (2019). Energy minimization for wireless communication with rotary-wing UAV. IEEE Transactions on Wireless Communications. DOI: https://doi.org/10.1109/TWC.2019.2902559.