IJEM Vol. 16, No. 4, 8 Aug. 2026
Cover page and Table of Contents: PDF (size: 2038KB)
PDF (2038KB), PP.356-370
Views: 0 Downloads: 0
Liver disease, bagging, boosting, feature importance, machine learning
The liver is one of the most essential internal organs in the human body, acting as a metabolic powerhouse and playing a key role in the immune system. However, Liver Diseases (LD) are rising globally, driven by unhealthy lifestyles and excessive alcohol use. Liver diseases cause millions of deaths annually all over the world, highlighting the need for early diagnosis. This study aims to evaluate the implication of ensemble machine learning techniques—bagging and boosting—for liver disease prediction, utilizing 30,691 instances and 11 features of the Liver Disease Patient Dataset (LDPD). To improve model performance, hyperparameter tuning, outlier removal, normalization for data scaling and feature importance to identify the most significant predictors are used. Eight state-of-the-art ensemble models were evaluated in this study: Random Forest (RF), Extra Trees Classifier (ETC), Bagged Decision Tree (Bagged DT), Adaptive Boosting (AdaBoost), Gradient Boosting (GradientBoost), Xtreme Gradient Boosting (XGBoost), Categorial Boosting (CatBoost) and Light Gradient Boosting Machine (LightGBM). Our experimental results showed that RF algorithm outperformed other algorithms, achieving the highest accuracy, specificity, precision, and F1-score of 99.85%, 99.85%, 99.93% and 99.89% respectively. While LightGBM attained the highest recall rate of 99.86% making it particularly suitable for identifying true positive cases and minimizing the missed diagnosis. These findings highlight the effectiveness of ensemble learning methods (bagging and boosting algorithms) in accurately predicting liver disease.
A. S. M. Shafi, "Performance Evaluation of Bagging and Boosting-Based Ensemble Learning Models for Clinical Liver Disease Prediction Using the LDPD Dataset", International Journal of Engineering and Manufacturing (IJEM), Vol.16, No.4, pp.356-370, 2026. DOI:10.5815/ijem.2026.04.24
[1]K. Sumeet, J.J. Larson, B. Yawn, T.M. Therneau, W.R. Kim (2013), Underestimation of liver-related mortality in the United States. Gastroenterology, 145:375–382, e371–372. https://doi.org/10.1053/j.gastro.2013.04.005.
[2]S.K. Asrani, H. Devarbhavi, J.Eaton, P.S. Kamath (2019), Burden of liver diseases in the world, J. Hepatol.70(1), 151–171. https://doi.org/10.1016/j.jhep.2018.09.014.
[3]S. Sontakke, J. Lohokare and R. Dani (2017), Diagnosis of liver diseases using machine learning, in International Conference on Emerging Trends & Innovation in ICT (ICEI), Pune, India, pp. 129–133. https://doi.org/10.1109/ETIICT.2017.7977023.
[4]A. Mohammed, R. Kora (2023), A comprehensive review on ensemble deep learning: opportunities and challenges, Journal of King Saud University - Computer and Information Sciences 35 (2), 757–774. https://doi.org/10.1016/j.jksuci.2023.01.014.
[5]O. Sagi, L. Rokach (2018), Ensemble learning: a survey, WIREs Data Mining and Knowledge Discovery 8 (4), e1249. https://doi.org/10.1002/widm.1249.
[6]C. Zhang, Y. Ma (Eds.) (2012), Ensemble Machine Learning: Methods and Applications, Springer, New York. https://doi.org/10.1007/978-1-4419-9326-7.
[7]Freund, Y., Schapire, R. E. (1996). Experiments with a new boosting algorithm. Proceedings of the Thirteenth International Conference on International Conference on Machine Learning, Vol. 96, pp. 148-156. https://dl.acm.org/doi/10.1145/3805689.3806539.
[8]Ghosh, Mounita & Raihan, Md. Mohsin & Raihan, M. & Akter, Laboni & Bairagi, Anupam & Alshamrani, Sultan & Masud, Mehedi. (2021). A Comparative Analysis of Machine Learning Algorithms to Predict Liver Disease. Intelligent Automation and Soft Computing. 29. 917-928. https://doi.org/10.32604/iasc.2021.017989.
[9]Islam, Rakibul & Sultana, Azrin & Tuhin, MD. (2024). A Comparative Analysis of Machine Learning Algorithms with Tree-Structured Parzen Estimator for Liver Disease Prediction. Healthcare Analytics. 6. 100358. https://doi.org/10.1016/j.health.2024.100358.
[10]Ganie, Shahid & Dutta Pramanik, Pijush. (2024). A comparative analysis of boosting algorithms for chronic liver disease prediction. Healthcare Analytics. 5. 100313. https://doi.org/10.1016/j.health.2024.100313.
[11]Ganie, Shahid & Dutta Pramanik, Pijush & Zhao, Zhongming. (2024). Improved liver disease prediction from clinical data through an evaluation of ensemble learning approaches. BMC Medical Informatics and Decision Making. 24. https://doi.org/10.1186/s12911-024-02550-y.
[12]Fırat ORHAN BULUCU, İrem ACER, Fatma LATİFOĞLU, Semra İÇER (2022). Predicting Liver Disease Using Decision Tree Ensemble Methods, Erciyes University, Journal of Institue of Science and Technology, Volume 38, Issue 2. https://izlik.org/JA77JX93LB.
[13]Deepthi, P., Gowtami, B., Vidyadhar, D., Shivanjaneya, T., & Vinay, V. (2024). Optimizing Liver Disease Prediction using SMOTE Integrated Supervised Learning Model. History of Medicine, 10(2), 20-30. DOI: 10.17720/2409-5834.v10.2.2024.03
[14]Amin, Ruhul, Rubia Yasmin, Sabba Ruhi, Md Habibur Rahman, and Md Shamim Reza (2023). Prediction of chronic liver disease patients using integrated projection based statistical feature extraction with machine learning algorithms, Informatics in Medicine Unlocked 36:101155. https://doi.org/10.1016/j.imu.2022.101155.
[15]Z. Cai, R.C. Poulos, J. Liu, Q. Zhong (2022). Machine learning for multi-omics data integration in cancer, iScience, 25 (2), Article 103798, https://doi.org/10.1016/j.isci.2022.103798.
[16]S. Sreejith, H. Khanna Nehemiah, A. Kannan (2020). Clinical data classification using an enhanced SMOTE and chaotic evolutionary feature selection, Computers in Biology and Medicine, Volume 126, 103991, ISSN 0010-4825, https://doi.org/10.1016/j.compbiomed.2020.103991.
[17]P. Kumar, R.S. Thakur (2021). Liver disorder detection using variable- neighbor weighted fuzzy K nearest neighbor approach, Multimedia Tools and Applications, 80 (11), pp. 16515-16535, https://doi.org/10.1007/s11042-019-07978-3.
[18]Dritsas, E., & Trigka, M. (2023). Supervised Machine Learning Models for Liver Disease Risk Prediction. Computers, 12(1), 19. https://doi.org/10.3390/computers12010019.
[19]Geurts, P., Ernst, D. & Wehenkel, L. (2006). Extremely randomized trees. Mach Learn 63, 3–42. https://doi.org/10.1007/s10994-006-6226-1.
[20]Rao, J.S., & Potts, W.J. (1997). Visualizing Bagged Decision Trees. Knowledge Discovery and Data Mining.
[21]Gong Cheng, Junwei Han (2016). A survey on object detection in optical remote sensing images, ISPRS Journal of Photogrammetry and Remote Sensing, Volume 117, Pages 11-28, ISSN 0924-2716, https://doi.org/10.1016/j.isprsjprs.2016.03.014.
[22]Dechun Zhao, Yi Wang, Qiangqiang Wang, Xing Wang (2019). Comparative analysis of different characteristics of automatic sleep stages, Computer Methods and Programs in Biomedicine, Volume 175, Pages 53-72, ISSN 0169-2607, https://doi.org/10.1016/j.cmpb.2019.04.004.
[23]Ananya Malik, Yash Tejas Javeri, Manav Shah, Ramchandra Mangrulkar (2022). Impact analysis of COVID-19 news headlines on global economy, Cyber-Physical Systems, Academic Press, Pages 189-206, ISBN 9780128245576, https://doi.org/10.1016/B978-0-12-824557-6.00001-7.
[24]Hoss Belyadi, Alireza Haghighat, Machine Learning Guide for Oil and Gas Using Python (2021), Gulf Professional Publishing, Pages 169-295, ISBN 9780128219294, https://doi.org/10.1016/B978-0-12-821929-4.00004-4.
[25]A. V. Dorogush, V. Ershov, and A. Gulin (2018). Catboost: Gradient boosting with categorical features support. arXiv: 1810.11363 [cs.LG]
[26]G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T.-Y. Liu (2017). Lightgbm: A highly efficient gradient boosting decision tree, Advances in neural information processing systems, pp. 3146–3154.
[27]Shahid Mohammad Ganie, Pijush Kanti Dutta Pramanik (2024), A comparative analysis of boosting algorithms for chronic liver disease prediction, Healthcare Analytics, Volume 5,100313, ISSN 2772-4425, https://doi.org/10.1016/j.health.2024.100313.
Chikodili, Nwodo & Mohammed, Abdulmalik & Abisoye, Opeyemi. (2021). Outlier Detection in Multivariate Time Series Data Using a Fusion of K-Medoid, Standardized Euclidean Distance and Z-Score. https://doi.org/10.1007/978-3-030-69143-1_21