IJWMT Vol. 16, No. 4, 8 Aug. 2026
Cover page and Table of Contents: PDF (size: 649KB)
PDF (649KB), PP.275-290
Views: 0 Downloads: 0
Decentralized Meta-Reinforcement Learning, Dynamic Spectrum Access, Cognitive Radio, 5G IoT, Spectrum Sharing, Multi-Agent System, Adaptive Resource Management
Dynamic spectrum access (DSA) in 5G IoT setups with cognitive radio is characterized by rapid and decentralized decision-making processes in highly non-stationary wireless environments, limited communication needs, and restrictive bounds. In this work, we present F-DMRL, a federated, communication-efficient decentralized meta-reinforcement learning framework for allowing a massive number of IoT devices to meta-learn collectively about spectrum-access strategies in a decentralized way without centralized control and without an extensive amount of inter-agent communication. Our method incorporates lightweight federated meta-parameter aggregation with gradient sparsification and periodic communication, allowing devices to only compress the meta-updates during this process and then adapt locally for task-specificity. We have presented analytical speedup guarantees and upper bounds on communication cost under bounded environmental drift and shown that using the approach proposed here, F-DMRL preserves convergence properties while posing a large reduction in coordination overhead at the same time. Simulations across various 5G IoT spectrum environments showed that F-DMRL performed faster adaptation (up to 45% fewer episodes), higher spectral efficiency, and lower interference probability compared to centralized meta-RL, federated DRL, and traditional decentralized RL baselines. Simulation results averaged across 10 independent runs demonstrate improvements of 45% faster adaptation and 60–80% lower communication overhead relative to baseline methods, while maintaining stable convergence.
Jayesh Kumar Dabi, Priyadarshi Ashok Dahat, "Federated and Communication-Efficient Decentralized Meta-Reinforcement Learning for Dynamic Spectrum Access in Cognitive Radio–Enabled 5G IoT Networks", International Journal of Wireless and Microwave Technologies(IJWMT), Vol.16, No.4, pp. 275-290, 2026. DOI:10.5815/ijwmt.2026.04.15
[1]Q. Zhao and B. M. Sadler, “A Survey of Dynamic Spectrum Access,” IEEE Signal Processing Magazine, Vol. 24, No. 3, pp. 79-89, 2007. doi:10.1109/MSP.2007.361604
[2]S. Haykin, “Cognitive radio: Brain-empowered wireless communications,” IEEE Journal on Selected Areas in Communications. vol. 23, no. 2, pp. 201-220, Feb. 2005, doi: 10.1109/JSAC.2004.839380.
[3]Ian F. Akyildiz., W.-Y. Lee., M. C. Vuran., & S. Mohanty, “NeXt generation/dynamic spectrum access/cognitive radio wireless networks: A Survey,” Computer Networks, vol. 50, no. 13, pp. 2127-2159, Sep. 2006, doi.org/10.1016/j.comnet.2006.05.001
[4]L. Chen., Z. Wang., X. Zhao., X. Shen & W. He, “A dynamic spectrum access algorithm based on deep reinforcement learning with novel multi-vehicle reward functions in cognitive vehicular networks,” Telecommunication Systems. vol. 87, Issue 2, Pages 359 – 383, June. 2024. https://doi.org/10.1007/s11235-024-01188-5
[5]C. Jiang., H. Zhang., Y. Ren., Z. Han., K.-C. Chen., & L. Hanzo, “Machine learning paradigms for next-generation wireless networks,” IEEE Wireless Communications. vol. 24, no. 2, pp. 98–105, Apr. 2017. doi: 10.1109/MWC.2016.1500356WC.
[6]Y. Zhou., F. Zhou., Y. Wu., R. Q. Hu., & Y. Wang. (2020), “Subcarrier assignment schemes based on Q-learning in wideband cognitive radio networks,” IEEE Transactions on Vehicular Technology. vol. 69, no. 1, pp. 1168-1172, Jan. 2020. doi: 10.1109/TVT.2019.2953809.
[7]J. Foerster., G. Farquhar., T. Afouras., N. Nardelli., & S. Whiteson, “Counterfactual multi-agent policy gradients,” Proceedings of the AAAI Conference on Artificial Intelligence. Article No.: 363, pp. 2974 – 2982, 2018. doi.org/10.1609/AAAI.V32I1.11794.
[8]L. Deng, M. Raissi & M. Xiao. (2026), “Meta-Learning-Based Surrogate Models for Efficient Hyperparameter Optimization,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 3, pp. 3931-3938, March 2026. doi: 10.1109/TPAMI.2025.3630178.
[9]S. Guo., Y. Du., & L. Liu. (2023), “A meta reinforcement learning approach for SFC placement in dynamic IoT-MEC networks,” Applied Sciences. vol. 13, no. 17, Art. no. 9960, Sep. 2023. doi: 10.3390/app13179960.
[10]S. Wang., H. Liu., P. H. Gomes., & B. Krishnamachari, “Deep reinforcement learning for dynamic multichannel access in wireless networks,” IEEE Transactions on Cognitive Communications and Networking. vol. 4, no. 2, Feb 2018. doi:10.1109/TCCN.2018.2809722.
[11]N. Zhao., H. Wu., F. R. Yu., L. Wang., W. Zhang., & V. C. M. Leung, “Deep-Reinforcement-Learning-Based Latency Minimization in Edge Intelligence over Vehicular Networks,” IEEE Internet of Things Journal. vol. 9, no. 2, pp. 1300–1312, Jan. 2022. doi: 10.1109/JIOT.2021.3078480.
[12]W. Y. B. Lim., N. C. Luong., D. T. Hoang., Y. Jiao., Y.-C. Liang., Q. Yang., D. Niyato., & C. Miao, “Federated learning in mobile edge networks: A comprehensive survey,” IEEE Communications Surveys & Tutorials. vol. 22, no. 3, pp. 2031–2063, Third Quarter 2020. doi: 10.1109/COMST.2020.2986024.
[13]M. Marwani. & G. Kaddoum, “Graph Neural Networks Approach for Joint Wireless Power Control and Spectrum Allocation,” IEEE Transactions on Machine Learning in Communications and Networking, vol. 2, pp. 717-732, June 2024. doi: 10.1109/TMLCN.2024.3408723.
[14]H. Song., L. Liu., J. Ashdown., & Y. Yi, “A deep reinforcement learning framework for spectrum management in dynamic spectrum access,” IEEE Internet of Things Journal, vol. 8, no. 14, pp. 11208-11218, 15 July 2021. doi: 10.1109/JIOT.2021.3052691.
[15]H.-H. Chang., Y. Song., T. T. Doan., & L. Liu, “Federated multi-agent deep reinforcement learning (Fed-MADRL) for dynamic spectrum access,” IEEE Transactions on Wireless Communications. vol. 22, no. 8, pp. 5337-5348, Aug. 2023. doi: 10.1109/TWC.2022.3233436.
[16]H. Qi., W. Ma., & B. Zhao, “Federated heterogeneous multi-agent deep reinforcement learning-based attack resilience scheduling for heterogeneous multi-integrated energy system,” Energy Reports, vol.14, pp.116-127, Dec. 2025. doi: 10.1016/j.egyr.2025.05.085.
[17]X. Tan., L. Zhou., H. Wang., Y. Sun., H. Zhao., B.-C. Seet., J. Wei., & V. C. M. Leung, “Cooperative multi-agent reinforcement learning based distributed dynamic spectrum access in cognitive radio networks,” IEEE Internet of Things Journal, vol. 9, no. 19, pp. 19477-19488, 1 Oct.1, 2022. doi: 10.1109/JIOT.2022. 3168296..
[18]H. Zhang., N. Liu., X. Chu., K. Long., A.-H. Aghvami., & V. C. M. Leung, “Network slicing based 5G and future mobile networks: Mobility, resource management, and challenges,” IEEE Communications Magazine, vol. 55, no. 8, pp. 138-145, Aug. 2017. doi: 10.1109/MCOM.2017.1600940.
[19]A. Al-Fuqaha., M. Guizani., M. Mohammadi., M. Aledhari., & M. Ayyash, “Internet of things: A survey on enabling technologies, protocols, and applications,” IEEE Communications Surveys & Tutorials. IEEE Communications Surveys & Tutorials, Vol. 17, Issue 4, pp. 2347 – 2376,2015. doi:10.1109/COMST.2015.2444095.
[20]Y. Zeng., J. Xu., & R. Zhang, “Energy minimization for wireless communication with rotary-wing UAV,” IEEE Transactions on Wireless Communications, vol. 18, no. 4, pp. 2329-2345, April 2019. doi:10.1109/TWC.2019.2902559.