Multi-Task BanglaBERT for Joint Sentiment and Fake News Detection in COVID-19 Discourse

PDF (760KB), PP.29-44

Views: 0 Downloads: 0

Author(s)

Arshadul Hoque 1

1. Department of Computer Science and Engineering, University of Chittagong, Chittagong, 4331, Bangladesh

* Corresponding author.

DOI: https://doi.org/10.5815/ijeme.2026.04.03

Received: 9 Jun. 2026 / Revised: 25 Jun. 2026 / Accepted: 13 Jul. 2026 / Published: 8 Aug. 2026

Index Terms

BanglaBERT, COVID-19 misinformation, Multi-Task learning, sentiment analysis, fake news detection, low-resource NLP

Abstract

The COVID-19 pandemic catalyzed an unprecedented surge of misinformation on social media, frequently intertwined with emotionally charged language. Understanding both the sentiment and truthfulness of this content is critical for public health monitoring and misinformation mitigation. However, Bangla—despite being a globally prominent language—remains severely underrepresented in joint sentiment and fake news detection research, with existing studies largely restricted to single-task settings. To bridge this gap, this paper proposes a novel multi-task BanglaBERT-based framework for the simultaneous classification of sentiment and truthfulness in COVID-19 discourse. Furthermore, we introduce the first publicly available, dual-annotated Bangla corpus for this domain, comprising 35,526 textual samples aggregated from social media and news sources. Our architecture employs a shared BanglaBERT encoder with dual task-specific heads, optimized using a task-prioritized loss function that combines modified Focal Loss and weighted cross-entropy to address inherent class imbalances. Extensive experiments demonstrate that the proposed model achieves 75.1% accuracy (Macro F1: 0.707) for sentiment classification and 88.0% accuracy (Macro F1: 0.851) for truthfulness detection. Ablation studies and error analyses confirm that our tailored loss strategies significantly enhance the recognition of underrepresented and semantically ambiguous classes, particularly neutral sentiments. By releasing our dataset, code, trained models, and a Gradio-based interactive demo, this work establishes a robust benchmark for multi-task learning in low-resource Bangla NLP and provides a practical tool for fact-checking during health crises.

Cite This Paper

Arshadul Hoque, “Multi-Task BanglaBERT for Joint Sentiment and Fake News Detection in COVID-19 Discourse”, International Journal of Education and Management Engineering (IJEME), Vol.16, No.4, pp. 16-28, 2026. DOI:10.5815/ijeme.2026.04.03

Reference

[1]Sabrina Jahan Maisha, Nuren nafisa, and Abdul Kadar Muhammad Masum. Supervised machine learning algorithms for sentiment analysis of bangla newspaper. International Journal of Innovative Computing, 2021. https://doi.org/10.11113/ijic.v11n2.321
[2]Md. Atikur Rahman and Emon Kumar Dey. Datasets for aspect-based sentiment analysis in bangla and its baseline evaluation. Data, 3(2), 2018. https://doi.org/10.3390/data3020015
[3]Salim Sazzed and Sampath Jayarathna. A sentiment classification in bengali and machine translated english corpus. In 2019 IEEE 20th International Conference on Information Reuse and Integration for Data Science (IRI), pages 107–114, 2019. https://doi.org/10.1109/iri.2019.00029
[4]Mohammad Nazmush Shamael, Sabila Nawshin, Swakkhar Shatabda, and Salekul Islam. Banglishrev: A large-scale bangla-english and code-mixed dataset of product reviews in e-commerce. arXiv preprint arXiv:2412.13161, 2024. https://doi.org/10.48550/arXiv.2412.13161
[5]Md. Zahin Hossain George, Naimul Hossain, Md. Rafiuzzaman Bhuiyan, Abu Kaisar Mohammad Masum, and Sheikh Abujar. Bangla fake news detection based on multichannel combined cnn-lstm. International Conference on Computing Communication and Networking Technologies, 2021. https://doi.org/10.2139/ssrn.5190043
[6]Md. Tanvir Rouf Shawon, G. M. Shahariar, Faisal Muhammad Shah, Mohammad Shafiul Alam, and Md. Shahriar Mahbub. Bengali fake review detection using semi-supervised generative adversarial networks. ICON, 2023. https://doi.org/10.1109/icnlp58431.2023.00011
[7]Fatema Tuj Johora Faria, Mukaffi Bin Moin, Zayeed Hasan, Md. Arafat Alam Khandaker, Niful Islam, Khan Md. Hasib, and M. F. Mridha. Multibanfakedetect: Integrating advanced fusion techniques for multimodal detection of bangla fake news in under-resourced contexts. Int. J. Inf. Manag. Data Insights, 2025. https://doi.org/10.1016/j.jjimei.2025.100347
[8]Zobaer Hossain, Ashraful Rahman, Saiful Islam, and Sudipta Kar. Banfakenews: A dataset for detecting fake news in bangla. arXiv: Computation and Language, 2020. https://doi.org/10.48550/arXiv.2004.08789
[9]Hrithik Majumdar Shibu, Shrestha Datta, Md. Sumon Miah, Nasrullah Sami, Mahruba Sharmin Chowd-hury, and Md Saiful Islam. From scarcity to capability: Empowering fake news detection in low-resource languages with LLMs. In Ruvan Weerasinghe, Isuri Anuradha, and Deshan Sumanathilaka, editors, Pro-ceedings of the First Workshop on Natural Language Processing for Indo-Aryan and Dravidian Languages, pages 100–107, Abu Dhabi, January 2025. Association for Computational Linguistics. https://doi.org/10.48550/arXiv.2501.09604
[10]Mohsinul Kabir, Obayed Bin Mahfuz, Hasan Mahmud, and Md. Kamrul Hasan. Banglabook: A large-scale bangla dataset for sentiment analysis from book reviews. Annual Meeting of the Association for Computational Linguistics, 2023. https://doi.org/10.18653/v1/2023.findings-acl.80
[11]Hemal Mahmud and Hasan Mahmud. Enhancing sentiment analysis in bengali texts: A hybrid approach using lexicon-based algorithm and pretrained language model bangla-bert. arXiv.org, 2024. https://doi.org/10.48550/arXiv.2411.19584
[12]Shihab Ahmed, Moythry Manir Samia, Maksuda Haider Sayma, M. Kabir, and M. F. Mridha. Trf-bert: A transformative approach to aspect-based sentiment analysis in the bengali language. PLoS ONE, 2024. https://doi.org/10.1371/journal.pone.0308050
[13]Shijie Chen, Yu Zhang, and Qiang Yang. Multi-task learning in natural language processing: An overview. ACM Computing Surveys, 2021. https://doi.org/10.48550/arXiv.2109.09138
[14]Sebastian Ruder. An overview of multi-task learning in deep neural networks. arXiv preprint arXiv:1706.05098, 2017. https://doi.org/10.48550/arXiv.1706.05098
[15]Lovre Torbarina, Tin Ferkovic, Łukasz Roguski, Velimir Mihelčić, Bruno Šarlija, and Željko Kraljević. Challenges and opportunities of using transformer-based multi-task learning in nlp through ml lifecycle: A survey. arXiv.org, 2023. https://doi.org/10.48550/arXiv.2308.08234
[16]Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. Multi-task deep neural networks for natural language understanding. Annual Meeting of the Association for Computational Linguistics, 2019. https://doi.org/10.48550/arXiv.1901.11504
[17]Yukang Xie, Chengyu Wang, Junbing Yan, Jiyong Zhou, Feiqi Deng, and Jun Huang. Making small language models better multi-task learners with mixture-of-task-adapters. Web Search and Data Mining, Pages 1094 – 1097, 2023. https://doi.org/10.1145/3616855.3635690
[18]Jonathan Pilault, Amine Elhattami, Christopher Pal, and Chris Pal. Conditionally adaptive multi-task learning: Improving transfer learning in nlp using fewer parameters & less data. arXiv: Learning, 2020. https://doi.org/10.48550/arXiv.2009.09139
[19]Abhik Bhattacharjee, Tahmid Hasan, Wasi Ahmad, Kazi Samin Mubasshir, Md Saiful Islam, Anindya Iqbal, M. Sohel Rahman, and Rifat Shahriyar. BanglaBERT: Language model pretraining and benchmarks for low-resource language understanding evaluation in Bangla. In Findings of the Association for Com-putational Linguistics: NAACL 2022, pages 1318–1327, Seattle, United States, July 2022. Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.findings-naacl.98
[20]Tahmid Hasan, Abhik Bhattacharjee, Kazi Samin, Masum Hasan, Madhusudan Basak, M. Sohel Rahman, and Rifat Shahriyar. Not low-resource anymore: Aligner ensembling, batch filtering, and new datasets for Bengali-English machine translation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2612–2623, Online, November 2020. Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.207
[21]Arshadul Hoque. Fine-tuned banglabert for bengali sentiment analysis. https://huggingface.co/ahs95/ banglabert-sentiment-analysis, 2025. DOI: Not available Enam Biswas. Bangla newspaper dataset - ebd, 2021. https://doi.org/10.34740/kaggle/dsv/3563095
[22]Md Ekramul Islam, Labib Chowdhury, Faisal Ahamed Khan, Shazzad Hossain, Md Sourave Hossain, Mohammad Mamun Or Rashid, Nabeel Mohammed, and Mohammad Ruhul Amin. Sentigold: A large bangla gold standard multi-domain sentiment analysis dataset and its evaluation. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 4207–4218, 2023. https://doi.org/10.1145/3580305.3599904
[23]Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(2):318–327, 2020. https://doi.org/10.1109/iccv.2017.324
[24]Fahmida Khanam, Anik Chakraborty, Md. Ahsan Habib, and Md. Sadiq Iqbal. Bangla sentiment analysis on highly imbalanced data using hybrid cnn-lstm & bangla bert. 2024 3rd International Conference on Advancement in Electrical and Electronic Engineering (ICAEEE), 2024. https://doi.org/10.1109/icaeee62219.2024.10561678
[25]Matteo Bodini and Matteo Bodini. Opinion mining from machine translated bangla reviews with stacked contractive auto-encoders. Journal of Ambient Intelligence and Humanized Computing, 2022. https://doi.org/10.1007/s12652-022-03760-w
[26]Sheikh Sadi Bandan, Md Sharuf Hossain., MD. Samiul Islam Sabbir, and Khadiza Tul Kobra. Deep learning for bengali fake news detection: Innovative approaches for accurate classification. International journal of research and innovation in applied science, 393-403, 2024. https://doi.org/10.51584/ijrias.2024.908034
[27]Md. Yasmi Tohabar, Nahiyan Nasrah, and Asif Mohammed Samir. Bengali fake news detection using machine learning and effectiveness of sentiment as a feature. International Conference on Informatics, Electronics and Vision, 2021. https://doi.org/10.1109/icievicivpr52578.2021.9564138
[28]S. Majumder and Dipankar Das. Studies towards language independent fake news detection. ICON, 439-446, 2021.
[29]Tanjina Sultana Camelia, Faizur Rahman Fahim, and M. Anwar. A regularized lstm method for detecting fake news articles. 2024 IEEE International Conference on Signal Processing, Information, Communication and Systems (SPICSCON), 2024. https://doi.org/10.1109/spicscon64195.2024.10941441
[30]Pronoy K. Mondal, Sadman Sadik Khan, Md. Masud Rana, Shahriar Sultan Ramit, A. Sattar, and Md. Sadekur Rahman. Breaking the fake news barrier: Deep learning approaches in bangla language. International Conference on Computing Communication and Networking Technologies, 2024. https://doi.org/10.1109/icccnt61001.2024.10725195
[31]Sheetal Harris, Hassan Jalil Hadi, Naveed Ahmad, and M. Alshara. Fake news detection revisited: An extensive review of theoretical frameworks, dataset assessments, model constraints, and forward-looking research agendas. Technologies, 2024. https://doi.org/10.3390/technologies12110222
[32]Quazi Adibur Rahman Adib, Md. Humaion Kabir Mehedi, Md. Sadman Sakib, Kabbya Kantam Patwary, Sabbir Hossain, and Annajiat Alim Rasel. A deep hybrid learning approach to detect bangla fake news. 2021 5th International Symposium on Multidisciplinary Studies and Innovative Technologies (ISMSIT), 2021. https://doi.org/10.1109/ismsit52890.2021.9604712
[33]Rich Caruana. Multitask learning. Machine learning, 28(1):41–75, 1997. https://doi.org/10.1007/978-1-4615-5529-2_5
[34]Jinlan Fu, See-Kiong Ng, and Pengfei Liu.  Polyglot prompt: Multilingual multitask prompt training. Conference on Empirical Methods in Natural Language Processing, 2022. https://doi.org/10.18653/v1/2022.emnlp-main.674
[35]Yiren Wang, ChengXiang Zhai, and Hany Hassan Awadalla. Multi-task learning for multilingual neural machine translation. arXiv: Computation and Language, 2020. https://doi.org/10.18653/v1/2020.emnlp-main.75
[36]Jing Lü, Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee. 12-in-1: Multi-task vision and language representation learning. arXiv: Computer Vision and Pattern Recognition, 2019. https://doi.org/10.48550/arXiv.1912.02315
[37]Chae-Gyun Lim, Young-Seob Jeong, and Ho-Jin Choi. Multi-task learning approach for utilizing temporal relations in natural language understanding tasks. Scientific Reports, 2023. https://doi.org/10.1038/s41598-023-35009-7
[38]Amirhossein Farzam, Shashank Shekhar, Isaac D. Mehlhaff, and Marco Morucci. Multi-task learning improves performance in deep argument mining models. Workshop on Argument Mining, 2023. https://doi.org/10.18653/v1/2024.argmining-1.5
[39]Clara Vania, Yova Kementchedjhieva, Anders Søgaard, and Adam Lopez. Edinburgh research explorer a systematic comparison of methods for low-resource dependency parsing on genuinely low-resource lan-guages. 1105–1116, 2019. https://doi.org/10.18653/v1/d19-1102
[40]Iker Garc’ia-Ferrero. Cross-lingual transfer for low-resource natural language processing. arXiv.org, 2025. https://doi.org/10.48550/arXiv.2502.02722
[41]Xiaobo Liang, Robert Mao, Lijun Wu, Juntao Li, Min Zhang, and Qing Li. Enhancing low-resource nlp by consistency training with data and model permutations. IEEE/ACM Transactions on Audio Speech and Language Processing, 189 – 199, 2024. https://doi.org/10.1109/taslp.2023.3325970
[42]Jeffrey Otoibhi, Oduguwa Damilola, and Okpare David. Sabiyarn: Advancing low resource languages with multitask nlp pretraining. Proceedings of the Sixth Workshop on African Natural Language Processing (AfricaNLP 2025), 95–107, 2025. https://doi.org/10.18653/v1/2025.africanlp-1.14
[43]Michael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, and Dietrich Klakow. A survey on recent approaches for natural language processing in low-resource scenarios. arXiv: Computation and Language, 2545–2568, 2020. https://doi.org/10.18653/v1/2021.naacl-main.201
[44]Emily Chang and Nada Basit. How many words does it take to understand a low-resource language? North American Chapter of the Association for Computational Linguistics, 207–224, 2025. https://doi.org/10.18653/v1/2025.naacl-srw.21
[45]Kumar Saunack, Saurav Kumar, and Pushpak Bhattacharyya. How low is too low? a monolingual take on lemmatisation in indian languages. North American Chapter of the Association for Computational Linguistics, 4088–4094, 2021. https://doi.org/10.18653/v1/2021.naacl-main.322
[46]Yudong Wang, Chao Ma, Qingxiu Dong, Lingpeng Kong, and Jun Xu. A challenging benchmark for low-resource learning. arXiv.org, 2057–2080, 2023. https://doi.org/10.18653/v1/2024.findings-acl.123
[47]Ping Yan and Jiangxu Wu. Low-resource named entity recognition based on multi-hop dependency trigger.
China National Conference on Chinese Computational Linguistics, 325–334, 2022. https://doi.org/10.1007/978-3-031-18315-7_21