IJCNIS Vol. 18, No. 4, 8 Aug. 2026
Cover page and Table of Contents: PDF (size: 1344KB)
PDF (1344KB), PP.169-186
Views: 0 Downloads: 0
Email Spam Detection, Ensemble Learning, Stacking Ensemble, Majority Voting Ensemble, NLP, Machine Learning Models
Spam emails continue to be a major issue in digital communication, hurting productivity and jeopardizing security. This paper presents a real-time spam email filtering method that combines ensemble learning techniques (stacking and majority voting) with machine learning classifiers such as Logistic Regression (LR), Support Vector Machine (SVM), Stochastic Gradient Descent (SGD), and Decision Trees (DT). The objective is to improve the accuracy and reliability of spam detection systems. The proposed models were tested against Kaggle’s Email Spam Dataset on crucial performance parameters such as accuracy, precision, recall, F-measure, false positive rate (FPR), and false negative rate (FNR). The results show that, while the stacking ensemble is competitive, the majority voting ensemble consistently achieved good results. It delivers improved accuracy of 98.59%, precision of 98.59%, recall of 98.59%, and F-measure of 98.59% while keeping lower FPR of 0.019 and FNR of 0.009, making it the most dependable option for real-world use. This study demonstrates the efficiency of ensemble learning, particularly majority voting, in building scalable and practical solutions for real-time spam email filtering. The findings lay a solid foundation for enhancing email security mechanisms and user experience.
Dharmaraj R. Patil, "Real-time Spam Email Filtering with Stacking and Majority Voting Ensembles", International Journal of Computer Network and Information Security(IJCNIS), Vol.18, No.4, pp. 169-186, 2026. DOI:10.5815/ijcnis.2026.04.09
[1]E. H. Tusher, M. A. Ismail, M. A. Rahman, A. H. Alenezi, and M. Uddin, “Email Spam: A Comprehensive Review of Optimize Detection Methods, Challenges, and Open Research Problems,” IEEE Access, 2024. DOI: 10.1109/AC- CESS.2024.3467996.
[2]E. G. Dada, J. S. Bassi, H. Chiroma, A. O. Adetunmbi, and O. E. Ajibuwa, “Machine learning for email spam filtering: review, approaches and open research problems,” Heliyon, vol. 5, no. 6, 2019. DOI: https://doi. org/10.1016/j.heliyon.2019.e01802.
[3]F. Ja´n˜ez-Martino, R. Alaiz-Rodr´ıguez, V. Gonza´lez-Castro, E. Fidalgo, and E. Alegre, “A review of spam email detection: analysis of spammer strategies and the dataset shift problem,” Artificial Intelligence Review, vol. 56, no. 2, pp. 1145-1173, 2023. DOI: https://doi.org/10.1007/s10462-022-10195-4.
[4]N. Ahmed, R. Amin, H. Aldabbas, D. Koundal, B. Alouffi, and T. Shah, “Machine learning techniques for spam detection in email and IoT platforms: analysis and research challenges,” Security and Communication Networks, vol. 2022, no. 1, Article ID 1862888, 2022. DOI: https://doi.org/10.1155/2022/1862888.
[5]P. Charanarur, H. Jain, G. S. Rao, D. Samanta, S. S. Sengar, and C. T. Hewage, “Machine-learning-based spam mail detector,” SN Computer Science, vol. 4, no. 6, Article ID 858, 2023. DOI: https://doi.org/10.1007/ s42979-023-02330-x.
[6]“Monthly share of spam in the total e-mail traffic worldwide from January 2014 to December 2023,” Available: https://www.statista.com/statistics/420391/spam-email-traffic-share/.
[7]EmailToolTester, ”EmailTooltester Deliverability Audit,” Available: https://www.emailtooltester. com/en/email-deliverability-audit/.
[8]Trend Micro, ”Email Threat Landscape Report: Cybercriminal Tactics, Techniques That Organizations Need to Know,” Available: https://www.trendmicro.com/ vinfo/us/security/research-and-analysis/threat-reports/roundup/annual-trend-micro-email-threats-report.
[9]C. N. Mohammed and A. M. Ahmed, “A semantic-based model with a hybrid feature engineering process for accurate spam detection,” Journal of Electrical Systems and Information Technology, vol. 11, no. 1, Article ID 26, 2024.
[10]Z. B. Siddique, M. A. Khan, I. U. Din, A. Almogren, I. Mohiuddin, and S. Nazir, “Machine learning-based detection of spam emails,” Scientific Programming, vol. 2021, no. 1, Article ID 6508784, 2021.
[11]A. J. Saleh, A. Karim, B. Shanmugam, S. Azam, K. Kannoorpatti, M. Jonkman, and F. De Boer, “An intelligent spam detection model based on artificial immune system,” Information, vol. 10, no. 6, Article ID 209, 2019.
[12]K. I. Roumeliotis, N. D. Tselikas, and D. K. Nasiopoulos, “Next-generation spam filtering: Comparative fine-tuning of LLMs, NLPs, and CNN models for email spam classification,” Electronics, vol. 13, no. 11, Article ID 2034, 2024.
[13]E. John-Africa and V. T. Emmah, “Performance evaluation of LSTM and RNN models in the detection of email spam messages,” European Journal of Information Technologies and Computer Science, vol. 2, no. 6, pp. 24-30, 2022.
[14]R. Indu and S. C. Dimri, “Detecting spam e-mails with content and weight-based binomial logistic model,” Journal of Web Engineering, vol. 22, no. 7, pp. 939-959, 2023.
[15]A. Sheneamer, “Comparison of deep and traditional learning methods for email spam filtering,” International Journal of Advanced Computer Science and Applications, vol. 12, no. 1, pp. 1-6, 2021.
[16]A. M. Bakare, K. S. M. Anbananthen, S. Muthaiyah, J. Krishnan, and S. Kannan, “Punctuation restoration with transformer model on social media data,” Applied Sciences, vol. 13, no. 3, p. 1685, 2023.
[17]G. Grefenstette, “Tokenization,” in Syntactic Wordclass Tagging, Dordrecht, Netherlands: Springer, 1999, pp. 117- 133.
[18]S. J. Mielke, Z. Alyafeai, E. Salesky, C. Raffel, M. Dey, M. Galle´, A. Raja, et al., “Between words and characters: A brief history of open-vocabulary modeling and tokenization in NLP,” arXiv preprint arXiv:2112.10508, 2021.
[19]S. Sarica and J. Luo, “Stopwords in technical language processing,” PLOS One, vol. 16, no. 8, p. e0254937, 2021.
[20]V. K. Pant, R. Sharma, and S. Kundu, “An overview of stemming and lemmatization techniques,” in Advances in Networks, Intelligence and Computing, pp. 308-321.
[21]T. Turki and S. S. Roy, “Novel hate speech detection using word cloud visualization and ensemble learning coupled with count vectorizer,” Applied Sciences, vol. 12, no. 13, p. 6611, 2022.
[22]G. M. Raza, Z. S. Butt, S. Latif, and A. Wahid, “Sentiment analysis on COVID tweets: An experimental analysis on the impact of count vectorizer and TF-IDF on sentiment predictions using deep learning models,” in Proc. 2021 Int. Conf. Digital Futures and Transformative Technologies (ICoDT2), pp. 1-6, 2021.
[23]M. M. Danyal, S. S. Khan, M. Khan, S. Ullah, M. B. Ghaffar, and W. Khan, “Sentiment analysis of movie reviews based on NB approaches using TF–IDF and count vectorizer,” Social Netw. Anal. Mining, vol. 14, no. 1, pp. 1-15, 2024.
[24]S. Wehnert, V. Sudhi, S. Dureja, L. Kutty, S. Shahania, and E. W. De Luca, “Legal norm retrieval with variations of the BERT model combined with TF-IDF vectorization,” in Proc. 18th Int. Conf. Artif. Intell. Law, pp. 285-294, 2021.
[25]T. G. Nick and K. M. Campbell, “Logistic regression,” Topics Biostatistics, pp. 273-301, 2007.
[26]J. R. Quinlan, “Induction of decision trees,” Machine Learn., vol. 1, pp. 81-106, 1986.
[27]S. Suthaharan and S. Suthaharan, “Support vector machine,” in Machine Learn. Models Algorithms Big Data Classif.: Thinking Examples Effective Learn., pp. 207-235, 2016.
[28]D. A. Pisner and D. M. Schnyer, “Support vector machine,” in Machine Learn., pp. 101-121, Academic Press, 2020.
[29]S. Amari, “Backpropagation and stochastic gradient descent method,” Neurocomputing, vol. 5, no. 4-5, pp. 185-196, 1993.
[30]L. Bottou, “Stochastic gradient descent tricks,” in Neural Networks: Tricks of the Trade: Second Edition, pp. 421- 436, Springer Berlin Heidelberg, 2012.
[31]F. Divina, A. Gilson, F. Gome´z-Vela, M. Garc´ıa Torres, and J. F. Torres, “Stacking ensemble learning for short-term electricity consumption forecasting,” Energies, vol. 11, no. 4, p. 949, 2018.
[32]S. Rajagopal, P. P. Kundapur, and K. S. Hareesha, “A stacking ensemble for network intrusion detection using heterogeneous datasets,” Security and Communication Networks, vol. 2020, no. 1, p. 4586875, 2020.
[33]A. Dogan and D. Birant, “A weighted majority voting ensemble approach for classification,” in 2019 4th International Conference on Computer Science and Engineering (UBMK), pp. 1-6, 2019.
[34]K. Raza, “Improving the prediction accuracy of heart disease with ensemble learning and majority voting rule,” in U-Healthcare Monitoring Systems, pp. 179-196, Academic Press, 2019.
[35]P. Singhvi, ”Spam Email Classification Dataset,” Available: https://www.kaggle.com/datasets/ purusinghvi/email-spam-classification-dataset, 2024.
[36]F. Zouak, O. El Beqqali, and J. Riffi, “BERT-GraphSAGE: hybrid approach to spam detection,” Journal of Big Data, vol. 12, no. 1, p. 128, 2025.