Deep Convolutional Auto encoders with Structured Latent Space for Fashion-MNIST Denoising and Classification

PDF (1593KB), PP.371-393

Views: 0 Downloads: 0

Author(s)

Pattapu Sravani 1 P. Satya srinivasa babu 1 Inturu Bhavani Siva Phanindra 1 Shaik Hasane Ahammad 1 Ramachandran Thandaiah Prabu 2 Ahmed Nabih Zaki Rashed 3,*

1. Department of ECE, Koneru Lakshmaiah Education Foundation, Vaddeswaram, Andhra Pradesh, India

2. Department of ECE, Saveetha School of Engineering, Saveetha Institute of Medical and Technical Sciences, SIMATS, Saveetha University, Chennai, Tamilnadu, India

3. Electronics and Electrical Communications Engineering Department, Faculty of Electronic Engineering, Menouf 32951, Menoufia University, Egypt

* Corresponding author.

DOI: https://doi.org/10.5815/ijem.2026.04.25

Received: 3 Apr. 2026 / Revised: 22 Apr. 2026 / Accepted: 8 Jul. 2026 / Published: 8 Aug. 2026

Index Terms

Convolutional auto encoders, structured latent space, fashion image analysis, denoising, semi-supervised learning

Abstract

Fashion image analysis has a lot of trouble working with noisy data, especially when there aren't any good training examples to use. This happens because standard auto encoders can't keep class separability when the noise level changes. This paper introduces Structured Auto encoders (SAE), which combine convolutional encoder-decoder architectures with geometric constraints in the latent space. This strategy uses greedy layer-wise pre-training, and then it fine-tunes the model using two loss functions: structured latent space regularization and reconstruction error. The encoder has convolutional layers with 3×3 filters and ReLU functions that only turn on when they are needed. During testing, it was tested on Fashion-MNIST with noise levels ranging from 20% to 70%. It was also tested on MNIST, DeepFashion2, and 3D human pose datasets. The results of the experiment show that SAE can classify 85.4% of the time with only five samples labeled for each group. This is much better than the standard auto encoders (68.4%) and the baseline support vector machine (SVM) methods (84.32%). With the best setup, which has 1,000 hidden nodes and a learning rate of 0.1, images with up to 70% noise can be reconstructed well With 80% classification accuracy on clean test data, the latent space structured method allows for the solution of basic issues that are related to introducing geometric constraints between the classes so that the class boundaries remain intact even in a noisy environment.The proposed method achieves a classification accuracy of 85.4% ± 1.2%under standard evaluation settings, with robustness validated under noise levels of up to 70%. All results are obtained using standard training–testing splits and are averaged over multiple experimental runs to ensure reliability and reproducibility. The proposed semi-supervised performance is enhanced by increasing the number of discriminative features through the use of a highly structured representation of the data, as opposed to using an unstructured representation, which results in performance enhancement. The author presents an original framework capable of simultaneously providing robust image classification and denoising capability through geometric constraints established in the framework, which enhances the learning of discriminative features through the use of limited amounts of labeled data. The implementation of this framework in the real-world fashion industry would facilitate the processing of fashion images with respect to noise resistance and the ability to annotate images quickly.

Cite This Paper

Pattapu Sravani, P. Satya srinivasa babu, Inturu Bhavani Siva Phanindra, Shaik Hasane Ahammad, Ramachandran Thandaiah Prabu, Ahmed Nabih Zaki Rashed, "Deep Convolutional Auto encoders with Structured Latent Space for Fashion-MNIST Denoising and Classification", International Journal of Engineering and Manufacturing (IJEM), Vol.16, No.4, pp.371-393, 2026. DOI:10.5815/ijem.2026.04.25

Reference

[1]Pintelas, Emmanuel, Ioannis E. Livieris, and Panagiotis E. Pintelas. "A convolutional autoencoder topology for classification in high-dimensional noisy image datasets." Sensors 21.22 (2021): 7731. https://doi.org/10.3390/s21227731
[2]An, Hyosun, et al. "Conceptual framework of hybrid style in fashion image datasets for machine learning." Fashion and Textiles 10.1 (2023): 18. https://doi.org/10.1186/s40691-023-00338-8
[3]Chang, Yeong-Hwa, and Ya-Ying Zhang. "Deep learning for clothing style recognition using YOLOv5." Micromachines 13.10 (2022): 1678. https://doi.org/10.3390/mi13101678
[4]Li, Ningyang, Zhaohui Wang, and Faouzi Alaya Cheikh. "Discriminating spectral–spatial feature extraction for hyperspectral image classification: A review." Sensors 24.10 (2024): 2987. https://doi.org/10.3390/s24102987
[5]Liang, Jun, et al. "Nonlinear Dynamic Process Monitoring Based on Discriminative Denoising Autoencoder and Canonical Variate Analysis." Actuators. Vol. 13. No. 11. MDPI, 2024. https://doi.org/10.3390/act13110440
[6]Wu, QuanLin, et al. "Denoising masked autoencoders help robust classification." arXiv preprint arXiv:2210.06983 (2022). 
https://doi.org/10.48550/arXiv.2210.06983
[7]Zhang, Qianjun, and Lei Zhang. "Convolutional adaptive denoising autoencoders for hierarchical feature extraction." Frontiers of Computer Science 12.6 (2018): 1140-1148. https://doi.org/10.1007/s11704-016-6107-0
[8]Berahmand, K., Daneshfar, F., Salehi, E.S. et al. Autoencoders and their applications in machine learning: a survey. Artif Intell Rev 57, 28 (2024). https://doi.org/10.1007/s10462-023-10662-6
[9]Jungo, Janosch, et al. "Representation learning for wearable-based applications in the case of missing data." arXiv preprint arXiv:2401.05437 (2024). 
https://doi.org/10.48550/arXiv.24
[10]Wang, Xingmei, et al. "Cela: Cost-efficient language model alignment for ctr prediction." arXiv preprint arXiv:2405.10596 (2024). 
https://doi.org/10.48550/arXiv.2405.10596
[11]Xing, Xin, et al. "Self-supervised learning application on covid-19 chest x-ray image classification using masked autoencoder." Bioengineering 10.8 (2023): 901.  https://doi.org/10.3390/bioengineering10080901
[12]Badrinarayanan, Vijay, Alex Kendall, and Roberto Cipolla. "Segnet: A deep convolutional encoder-decoder architecture for image segmentation." IEEE transactions on pattern analysis and machine intelligence 39.12 (2017): 2481-2495. doi: 10.1109/TPAMI.2016.2644615
[13]Berahmand, Kamal, et al. "Autoencoders and their applications in machine learning: a survey." Artificial intelligence review 57.2 (2024): 28. doi.org/10.1007/s10462-023-10662-6
[14]Walczyna, Tomasz, Damian Jankowski, and Zbigniew Piotrowski. "Enhancing anomaly detection through latent space manipulation in autoencoders: A comparative analysis." Applied Sciences 15.1 (2024): 286. https://doi.org/10.3390/app15010286
[15]Lazebnik, Teddy, and Liron Simon-Keren. "Knowledge-integrated autoencoder model." Expert Systems with Applications 252 (2024): 124108. https://doi.org/10.1016/j.eswa.2024.124108
[16]Radha, K., and Yepuganti Karuna. "Latent space autoencoder generative adversarial model for retinal image synthesis and vessel segmentation." BMC Medical Imaging 25.1 (2025): 149. https://doi.org/10.1186/s12880-025-01694-1
[17]Cao, B., Bi, Z., Hu, Q. et al. AutoEncoder-Driven Multimodal Collaborative Learning for Medical Image Synthesis. Int J Comput Vis 131, 1995–2014 (2023). https://doi.org/10.1007/s11263-023-01791-0
[18]Liu, Jia, et al. "A feature selection method based on multiple feature subsets extraction and result fusion for improving classification performance." Applied Soft Computing 150 (2024): 111018. https://doi.org/10.1016/j.asoc.2023.111018
[19]de Oliveira, Emerson Vilar, Dunfrey Pires Aragão, and Luiz Marcos Garcia Gonçalves. "Elsa: Expanded latent space autoencoder for image feature extraction and classification." Proceedings Copyright 703 (2024): 710. https://doi.org/10.5220/0012455300003660
[20]Zhang, Junping, et al. "Recent advances in artificial intelligence generated content." Frontiers of Information Technology & Electronic Engineering 25.1 (2024): 1-5. https://doi.org/10.1631/FITEE.2410000
[21]Zhu, Yixuan, et al. "Stableswap: Stable face swapping in a shared and controllable latent space." IEEE Transactions on Multimedia 26 (2024): 7594-7607. https://doi.org/10.1109/TMM.2024.3369853.
[22]Luo, Zhongtang. "ICLR Points: How Many ICLR Publications Is One Paper in Each Area?." arXiv preprint arXiv:2503.16623 (2025). https://doi.org/10.48550/arXiv.2503.16623
[23]Peng, Xi, et al. "Structured auto encoders for subspace clustering." IEEE Transactions on Image Processing 27.10 (2018): 5076-5086. https://doi.org/10.1109/TIP.2018.2848470.
[24]Shajini, Majuran, and Amirthalingam Ramanan. "A knowledge-sharing semi-supervised approach for fashion clothes classification and attribute prediction." The Visual Computer 38.11 (2022): 3551-3561. https://doi.org/10.1007/s00371-021-02178-3.
[25]Thwe, Yamin, Nipat Jongsawat, and Anucha Tungkasthan. "A semi-supervised learning approach for automatic detection and fashion product category prediction with small training dataset using FC-YOLOv4." Applied Sciences 12.16 (2022): 8068.  https://doi.org/10.3390/app12168068
[26]Mukhamediev, Ravil I. "State-of-the-art results with the fashion-MNIST dataset." Mathematics 12.20 (2024): 3174. https://doi.org/10.3390/math12203174
[27]Sigdel, Madhav, et al. "Evaluation of semi-supervised learning for classification of protein crystallization imagery." IEEE SOUTHEASTCON 2014. IEEE, 2014. https://doi.org/10.1109/SECON.2014.6950649.
[28]Fan, Cheng, et al. "Integrating active learning and semi-supervised learning for improved data-driven HVAC fault diagnosis performance." Applied Energy 356 (2024): 122356. https://doi.org/10.1016/j.apenergy.2023.122356
[29]Roda, Hezi, and Amir B. Geva. "Semi-supervised active learning using convolutional auto-encoder and contrastive learning." Frontiers in Artificial Intelligence 7 (2024): 1398844.  https://doi.org/10.3389/frai.2024.1398844
[30]Rizve, Mamshad Nayeem, et al. "Openldn: Learning to discover novel classes for open-world semi-supervised learning." European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022. https://doi.org/10.1109/IMCOM64595.2025.10857517.
[31]Y. Salim, A. F. Rakasyah, H. Darwis, Harlinda, Irawati and A. R. Manga, "Comparative Analysis of Machine Learning Algorithms and Ensemble Techniques for Diverse Image Classification Tasks," 2025 19th International Conference on Ubiquitous Information Management and Communication (IMCOM), Bangkok, Thailand, 2025, pp. 1-7, https://doi.org/10.1109/IMCOM64595.2025.10857517.