CPU-Centric Evaluation of Lightweight CNNs for Four-Class Chest X-Ray Classification: Performance, Latency, and Quantization Limits

PDF (993KB), PP.330-345

Views: 0 Downloads: 0

Author(s)

Kateryna P. Hazdiuk 1,* Roman V. Movcheniuk 1

1. Department of Computer Systems Software, Yuriy Fedkovych Chernivtsi National University, Chernivtsi, 58012, Ukraine

* Corresponding author.

DOI: https://doi.org/10.5815/ijem.2026.05.18

Received: 5 Jun. 2026 / Revised: 17 Jul. 2026 / Accepted: 5 Sep. 2026 / Published: 8 Oct. 2026

Index Terms

Software engineering, deep learning, convolutional neural networks, chest X-ray, lightweight models, quantization, computational complexity, inference latency, Pareto analysis, medical diagnostics

Abstract

In most studies applying deep learning to medical image analysis, classification accuracy is the sole optimization criterion, while the computational cost of the resulting models remains undocumented: parameter and operation counts are reported only occasionally, and inference latency on a central processing unit is almost never published. This makes an informed model choice impossible for institutions without graphics accelerators, such as district hospitals, mobile diagnostic units and field hospitals. This paper experimentally investigates the trade-off between diagnostic performance and computational resources for four-class chest X-ray classification. The computational cost of 15 widely used architectures was evaluated with a single tool at an input resolution of 224×224, and the gap between the heaviest and the lightest architecture reaches a factor of 323. Four models – a heavy baseline, two lightweight architectures, and a custom compact network – were trained under a unified protocol on the COVID-19 Radiography Database (21,165 images across four classes) and profiled strictly on a CPU with a batch size of one. The three standard architectures were initialized with ImageNet-pretrained weights, and the custom network was trained from scratch. Each architecture was trained across three independent random seeds on a fixed data split, so all performance metrics are reported as mean ± std, and every comparison is accompanied by a confidence interval and a significance test. The lightweight architecture trailed the baseline by only 1.24 percentage points in macro-F1 (95% CI: 0.84 to 1.63, p = 0.005) while requiring 69 times fewer operations, achieving 19.3 times lower latency, and reducing model size by a factor of 15.4; the efficiency metric, defined as macro-F1 per GFLOP, differs by a factor of 68. Additionally, post-training quantization systematically failed across all three runs for both architectures combining depthwise convolutions with Squeeze-and-Excitation blocks, whereas the same pipeline quantized the remaining models without statistically significant loss. Neither depthwise convolutions nor the Hard-Swish activation accounted for the failure; instead, the number of channel-wise multiplication nodes in the exported graph separated the two groups exactly, making the risk identifiable directly from the static graph prior to training. The practical implication is that the quantizability of a lightweight model must not be assumed; it must be verified empirically.

Cite This Paper

Kateryna P. Hazdiuk, Roman V. Movcheniuk, "CPU-Centric Evaluation of Lightweight CNNs for Four-Class Chest X-Ray Classification: Performance, Latency, and Quantization Limits", International Journal of Engineering and Manufacturing (IJEM), Vol.16, No.5, pp. 330-345, 2026. DOI:10.5815/ijem.2026.05.18

Reference

[1]G. Litjens, T. Kooi, B. E. Bejnordi, A. A. A. Setio, et al., "A survey on deep learning in medical image analysis," Medical Image Analysis, vol. 42, pp. 60-88, 2017. doi:10.1016/j.media.2017.07.005
[2]P. Rajpurkar, J. Irvin, K. Zhu, et al., "CheXNet: radiologist-level pneumonia detection on chest X-rays with deep learning," arXiv preprint arXiv:1711.05225, 2017. doi:10.48550/arXiv.1711.05225
[3]M. E. H. Chowdhury, T. Rahman, A. Khandakar, et al., "Can AI help in screening viral and COVID-19 pneumonia?," IEEE Access, vol. 8, pp. 132665-132676, 2020. doi:10.1109/ACCESS.2020.3010287
[4]T. Rahman, A. Khandakar, Y. Qiblawey, et al., "Exploring the effect of image enhancement techniques on COVID-19 detection using chest X-ray images," Computers in Biology and Medicine, vol. 132, art. no. 104319, 2021. doi:10.1016/j.compbiomed.2021.104319
[5]D. S. Kermany, M. Goldbaum, W. Cai, et al., "Identifying medical diagnoses and treatable diseases by image-based deep learning," Cell, vol. 172, no. 5, pp. 1122-1131, 2018. doi:10.1016/j.cell.2018.02.010
[6]A. Howard, M. Sandler, G. Chu, et al., "Searching for MobileNetV3," in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 1314-1324. doi:10.1109/ICCV.2019.00140
[7]M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, "MobileNetV2: inverted residuals and linear bottlenecks," in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 4510-4520. doi:10.1109/CVPR.2018.00474
[8]M. Tan and Q. V. Le, "EfficientNet: rethinking model scaling for convolutional neural networks," in Proc. 36th International Conference on Machine Learning (ICML), PMLR, vol. 97, 2019, pp. 6105-6114. doi:10.48550/arXiv.1905.11946
[9]B. Jacob, S. Kligys, B. Chen, et al., "Quantization and training of neural networks for efficient integer-arithmetic-only inference," in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 2704-2713. doi:10.1109/CVPR.2018.00286
[10]R. Krishnamoorthi, "Quantizing deep convolutional networks for efficient inference: a whitepaper," arXiv preprint arXiv:1806.08342, 2018. doi:10.48550/arXiv.1806.08342
[11]M. Nagel, M. Fournarakis, R. A. Amjad, Y. Bondarenko, M. van Baalen, and T. Blankevoort, "A white paper on neural network quantization," arXiv preprint arXiv:2106.08295, 2021. doi:10.48550/arXiv.2106.08295
[12]C.-T. Yen and C.-Y. Tsao, "Lightweight convolutional neural network for chest X-ray images classification," Scientific Reports, vol. 14, no. 1, art. no. 29759, 2024. doi:10.1038/s41598-024-80826-z
[13]T. Shanableh, "LungVision: X-ray imagery classification for on-edge diagnosis applications," Algorithms, vol. 17, no. 7, art. no. 280, 2024. doi:10.3390/a17070280
[14]K. Sun, X. Wang, X. Miao, and Q. Zhao, "A review of AI edge devices and lightweight CNN and LLM deployment," Neurocomputing, vol. 614, art. no. 128791, 2025. doi:10.1016/j.neucom.2024.128791
[15]Z. Yuan, J. Liu, J. Wu, D. Yang, et al., "Benchmarking the reliability of post-training quantization: a particular focus on worst-case performance," in Proc. 2nd Workshop on New Frontiers in Adversarial Machine Learning (AdvML-Frontiers), ICML, Honolulu, HI, USA, 2023. [Online]. Available: https://openreview.net/forum?id=iKBdmOVJ5T (accessed on September 2026)
[16]S. Yoon, N. Kim, and H. Kim, "LAMP-Q: layer sensitivity-aware mixed-precision quantization for MobileNetV3," in Proc. 2025 International Conference on Electronics, Information, and Communication (ICEIC), Osaka, Japan, 2025, pp. 1-3. doi:10.1109/ICEIC64972.2025.10879604
[17]B. Rokh, A. Azarpeyvand, and A. Khanteymoori, "A comprehensive survey on model quantization for deep neural networks in image classification," ACM Transactions on Intelligent Systems and Technology, vol. 14, no. 6, pp. 1-50, 2023. doi:10.1145/3623402
[18]V. Sovrasov, "ptflops: a flops counting tool for neural networks in PyTorch framework," version 0.7.5, 2025. [Online]. Available: https://github.com/sovrasov/flops-counter.pytorch (accessed on August 2026)
[19]Microsoft, "ONNX Runtime: cross-platform accelerated machine learning," version 1.29.0, 2025. [Online]. Available: https://onnxruntime.ai (accessed on August 2026)
[20]K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition," in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770-778. doi:10.1109/CVPR.2016.90
[21]T. Sheng, C. Feng, S. Zhuo, X. Zhang, L. Shen, and M. Aleksic, "A quantization-friendly separable convolution for MobileNets," arXiv preprint arXiv:1803.08607, 2018. doi:10.48550/arXiv.1803.08607
[22]R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, "Grad-CAM: visual explanations from deep networks via gradient-based localization," International Journal of Computer Vision, vol. 128, no. 2, pp. 336-359, 2020. doi:10.1007/s11263-019-01228-7
[23]K. P. Hazdiuk and R. V. Movcheniuk, “Architecture of an X-ray image analysis system using GCP Vertex AI, Kubernetes, Dataflow and microservice architecture,” Vcheni zapysky TNU imeni V.I. Vernadskoho. Seriia: Tekhnichni nauky, 2026, p. 43-50 (in Ukrainian; cited here only for the cloud-architecture background of the prototype service of Section 5). doi:10.32782/2663-5941/2026.1.2/06