T. Haripriya

Work place: Department of Mathematics, Sreyas Institute of Engineering and Technology, Hyderabad, 500068, Telangana, India

E-mail: haripriya.tirumala@gmail.com

Website: https://orcid.org/0009-0007-3938-2703

Research Interests:

Biography

Dr. T. Hari Priya is a Professor in the Department of Mathematics at Sreyas Institute of Engineering and Technology. She holds a Ph.D. in Mathematics and has over 25 years of teaching experience in higher education. Her academic interests include mathematical sciences, research, and higher education. She has published several quality research papers in reputed national and international journals and has contributed significantly to teaching, student mentoring, and academic development. Deep learning models, explainable AI, and advanced machine learning techniques for real-world applications.

Author Articles
An Interpretable Machine Learning Framework for Breast Cancer Diagnosis Using Statistical Feature Analysis and Ensemble Classification

By T. Haripriya M.V. Ramana Murthy Ch. Vasavi Swathi Gowroju G. Srinivas Devineni Gireesh Kumar

DOI: https://doi.org/10.5815/ijem.2026.04.16, Pub. Date: 8 Aug. 2026

A clear, statistically sound, yet easily understandable breast cancer diagnosis is a difficult issue in all healthcare systems, because early stages of breast cancer are critical in therapy success and long-term survivability. This machine-learning-based breast cancer classifier, in a statistically justified, rigorously experimentally validated way, classifies a set of 569 breast cancer cases with 9 cytological features for breast cancer diagnosis. The classifier uses a rigorous set of data cleanup measures, including missing-value substitution, correlation-based feature reduction, and projection into principal component space, to achieve high data quality, reduce redundancy, and enhance feature usefulness. Five supervised classifiers, in a widely accepted train-test model using an 80:20 random sample split and 5-fold cross-validation, are fitted and evaluated using Accuracy, Precision, Recall, F1-score, and Area under the ROC curve. In these tests, the Random Forest classifier got the best result, with 95.84% Accuracy, 95.31% Precision, 95.12% Recall, 95.21% F1-score and 0.982 area under the ROC curve; in a statistically sound consistency test using cross-validation, its mean accuracy reached 95.96% with a small standard deviation of 0.43. To provide a clear, interpretable indication of which features truly matter, we performed a feature-importance analysis on the best classifier, the Random Forest model. Results show that the expression levels of Bland Chromatin, Single Epithelial Cell Size, Normal Nucleoli, Uniformity of Cell Shape, Uniformity of Cell Size and Bare Nuclei are closely related to breast cancer diagnosis; this is almost the same as the clinical diagnosis findings, and very naturally suggests that abnormalities of cellular morphology and nuclei are major symptoms of breast cancer. In comparison, prior research may neglect validation and efficiency comparisons or focus only on the classifier's accuracy. Our method combines multiple levels of assessment (statistical data-by-data validation, feature importance, cross-validation, and comparison of different classifiers using ensemble learning) into a single evaluation system. This combined approach not only enhances predictive capability but also makes the entire setup more explicitly interpretable from a clinical perspective, thereby making it more suitable for health care decision support. Given the strong classification performance, interpretability, and validation suggested above, the model would help physicians detect breast cancer very early, reducing the risk of misdiagnosis.

[...] Read more.
Other Articles