Work place: Department of Electronics and Communication Engineering, Koneru Lakshmaiah Education Foundation, Vaddeswaram, Guntur District, India
E-mail: krishnamohan506@gmail.com
Website: https://orcid.org/0009-0007-6454-8090
Research Interests:
Biography
Dr. Kaja Krishna Mohan obtained his B.Tech and M.Tech degrees from JNTUK, Kakinada. He obtained his Ph.D from K L University. Now he is working as Assistant Professor in the Department of Electronics and Communication Engineering. He has published several papers in reputed journals. His research interest includes Image processing, Image classification, Deep learning and Ensemble learning.
By Mudunuru Suneel Kaja Krishna Mohan Banothu Yedukondala Venkata Naga Raja Swamy Seva Sreedhar Babu Venkata Raghavendra Miriampally P. Rama Koteswara Rao Kama Ramudu
DOI: https://doi.org/10.5815/ijwmt.2026.05.24, Pub. Date: 8 Oct. 2026
Face recognition systems are increasingly deployed in security-critical applications, but remain vulnerable to presentation attacks such as printed photographs, replay videos, and three-dimensional masks. Although recent face anti-spoofing methods exploit complementary information from RGB, depth, and near-infrared (NIR) representations, existing approaches primarily focus on feature fusion and do not explicitly model the intrinsic consistency relationships among
these representations. This limitation can reduce robustness against sophisticated spoofing attacks and cross-dataset variations.
This paper proposes a novel Cross-Modal Consistency Learning (CMCL) framework for robust face anti-spoofing using the original RGB input together with RGB-derived depth and NIR representations. The depth and NIR representations are constructed from RGB inputs during the modality preprocessing stage and subsequently processed together with RGB using independent Swin Transformer encoders. A cross-modal attention fusion module adaptively integrates complementary appearance, geometric, and spectral information, while a consistency learning module explicitly encourages feature coherence for genuine samples and emphasizes cross-modal discrepancies associated with spoof attacks. The consistency module is used only during training and removed during inference, avoiding additional deployment overhead.
Extensive experiments are conducted on CASIA-SURF, WMCA, CelebA-Spoof, and MSU-MFSD using intra-dataset, cross-dataset, ablation, feature-space, modality-consistency, and statistical analyses. The proposed CMCL achieves an ACER of 0.70% and accuracy of 99.3% on CASIA-SURF, while achieving an ACER of 2.55% and accuracy of 97.9% on WMCA. The results demonstrate improved spoof detection performance, representation discrimination, and cross-dataset generalization compared with the evaluated state-of-the-art methods. The proposed framework provides a consistency-driven and computationally practical approach for robust face anti-spoofing.
Subscribe to receive issue release notifications and newsletters from MECS Press journals