Multilingual Fake News Detection Using Machine Learning with Contextual-Based Feature Extraction

PDF (556KB), PP.97-105

Views: 0 Downloads: 0

Author(s)

Nikita Garg 1,* Pritam Singh Negi 1

1. Department of Computer Science & Engineering, HNB Garhwal University (A Central University), Srinagar Garhwal- 246 174, Uttarakhand, India

* Corresponding author.

DOI: https://doi.org/10.5815/ijem.2026.04.07

Received: 5 Mar. 2026 / Revised: 12 Apr. 2026 / Accepted: 16 Jul. 2026 / Published: 8 Aug. 2026

Index Terms

Fake News Detection, Multilingual Natural Language Processing, Feature Extraction, Contextual Embeddings, N-gram, Machine Learning

Abstract

Fake news has become a major challenge in today’s digital environment, particularly in languages where labeled data is limited. Most existing research has primarily focused on English due to the easy availability of annotated datasets, whereas low-resource languages such as Bengali remain underexplored. This study presents a multilingual approach for fake news detection using machine learning with contextual-based feature extraction. The proposed method integrates n-gram techniques with sentence-level contextual embeddings to capture both word-level patterns and semantic meaning. Since labeled data is not available for the Bengali dataset, a translation-based strategy is employed, followed by a pseudo-labeling process to assign labels automatically. The models are trained on English news titles and subsequently evaluated on both English and Bengali datasets to examine their cross-lingual effectiveness. The experimental findings indicate that ensemble-based classifiers such as Random Forest and Gradient Boosting achieve reliable performance across both languages. In some cases, the results for Bengali data are comparable or slightly better than those for English. The study demonstrates that effective fake news detection is possible in low-resource languages using short text data without relying on manually labeled datasets. The proposed approach provides a simple and efficient solution for multilingual fake news detection in data-scarce environments.

Cite This Paper

Nikita Garg, Pritam Singh Negi, "Multilingual Fake News Detection Using Machine Learning with Contextual-Based Feature Extraction", International Journal of Engineering and Manufacturing (IJEM), Vol.16, No.4, pp.97-105, 2026. DOI:10.5815/ijem.2026.04.07

Reference

[1]S. Vosoughi, D. Roy, and S. Aral, “The spread of true and false news online,” Science, vol. 359, no. 6380, pp. 1146–1151, 2018. https://doi.org/10.1126/science.aap9559.
[2]K. Shu, A. Sliva, S. Wang, J. Tang, and H. Liu, “Fake news detection on social media: A data mining perspective,” SIGKDD Explorations, vol. 19, no. 1, pp. 22–36, 2017. https://doi.org/10.1145/3137597.3137600.
[3]D.-H. Lee, “Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks,” ICML Workshop, 2013. https://arxiv.org/abs/1303.0745.
[4] S. Ruder, I. Vulić, and A. Søgaard, “A survey of cross-lingual word embedding models,” Journal of Artificial Intelligence Research, vol. 65, pp. 569–631, 2019. https://doi.org/10.1613/jair.1.11640.
[5]J. H. Lau and T. Baldwin, “An empirical evaluation of doc2vec with practical insights into document embedding generation,” ACL Workshop, 2016. https://aclanthology.org/W16-1609/.
[6]V. Pérez-Rosas, B. Kleinberg, A. Lefevre, and R. Mihalcea, “Automatic detection of fake news,” COLING, 2018. https://doi.org/10.18653/v1/C18-1287.
[7]J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” NAACL, 2019. https://doi.org/10.18653/v1/N19-1423.
[8]S. Pan and Q. Yang, “A survey on transfer learning,” IEEE Transactions on Knowledge and Data Engineering, vol. 22, no. 10, pp. 1345–1359, 2010. https://doi.org/10.1109/TKDE.2009.191.
[9]A. Conneau et al., “Unsupervised cross-lingual representation learning,” ACL, 2018. https://doi.org/10.18653/v1/P18-1073.
[10]A. Conneau et al., “Cross-lingual language model pretraining,” NeurIPS, 2020. https://arxiv.org/abs/1911.02116.
[11]L. Breiman, “Random forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001. https://doi.org/10.1023/A:1010933404324.
[12]K. Sohn et al., “FixMatch: Simplifying semi-supervised learning with consistency and confidence,” NeurIPS, 2020. https://arxiv.org/abs/2001.07685.
[13]K. K. Yadav, G. Thakur, and J. Srivastava, “A Hybrid Tri-Encoder model for fake news detection in Bengali with LIME-based explainability,” Intelligent Decision Technologies, vol. 19, no. 1, 2025. https://doi.org/10.1177/18724981251384403.
[14]A. Bhosle et al., “Multilingual Fake News Detection: A Machine Learning Approach for Indian Languages,” Proceedings of ICSIAIML, 2025. https://doi.org/10.2991/978-94-6463-948-3_28.
[15]N. Garg and P. S. Negi, “Context-Aware Multilingual Fake News Detection Using Machine Learning and Genetic Algorithm-Based Feature Selection,” International Journal for Research in Applied Science & Engineering Technology (IJRASET), vol. 13, no. 11, 2025. https://doi.org/10.22214/ijraset.2025.75362.
[16]M. Hossain et al., “Transformers for Bengali Misinformation Detection,” Journal of Natural Language Processing, vol. 32, no. 2, 2025. https://doi.org/10.3389/frai.2025.1537432.
[17]A. Khan et al., “Multilingual Fake News Detection in Low-Resource Languages: A 2024 Survey,” IEEE Access, vol. 12, 2024. https://doi.org/10.1016/j.knosys.2024.111884.