Cross-Modal Consistency Learning for Robust Face Anti-Spoofing Using RGB, RGB-Derived Depth and NIR Representations

PDF (1508KB), PP.398-427

Views: 0 Downloads: 0

Author(s)

Mudunuru Suneel 1,* Kaja Krishna Mohan 2 Banothu Yedukondala Venkata Naga Raja Swamy 3 Seva Sreedhar Babu 4 Venkata Raghavendra Miriampally 5 P. Rama Koteswara Rao 6 Kama Ramudu 1

1. Department of Electronics and Communication Engineering, Aditya University, Surampalem, Andhra Pradesh, India

2. Department of Electronics and Communication Engineering, Koneru Lakshmaiah Education Foundation, Vaddeswaram, Guntur District, India

3. Department of Electronics and Communication Engineering, Lakireddy Bali Reddy College of Engineering, Mylavaram, Andhra Pradesh, India

4. Department of Electronics and Communication Engineering, Sree Vahini Institute of Science and Technology, Thiruvuru, Andhra Pradesh, India

5. Department of Electronics and Communication Engineering, Sree Dattha Group of Institutions, Sheriguda, Hyderabad, Telangana, India

6. Department of Computer Science and Engineering, Sree Dattha Institute of Engineering & Science, Sheriguda, Hyderabad, Telangana, India

* Corresponding author.

DOI: https://doi.org/10.5815/ijwmt.2026.05.24

Received: 22 Jul. 2026 / Revised: 17 Aug. 2026 / Accepted: 8 Sep. 2026 / Published: 8 Oct. 2026

Index Terms

Face Anti-Spoofing, Cross-Modal Consistency Learning, Swin Transformer, Cross-Modal Attention, Multi-Modal Feature Fusion, Biometric Security

Abstract

Face recognition systems are increasingly deployed in security-critical applications, but remain vulnerable to presentation attacks such as printed photographs, replay videos, and three-dimensional masks. Although recent face anti-spoofing methods exploit complementary information from RGB, depth, and near-infrared (NIR) representations, existing approaches primarily focus on feature fusion and do not explicitly model the intrinsic consistency relationships among 
these representations. This limitation can reduce robustness against sophisticated spoofing attacks and cross-dataset variations.
This paper proposes a novel Cross-Modal Consistency Learning (CMCL) framework for robust face anti-spoofing using the original RGB input together with RGB-derived depth and NIR representations. The depth and NIR representations are constructed from RGB inputs during the modality preprocessing stage and subsequently processed together with RGB using independent Swin Transformer encoders. A cross-modal attention fusion module adaptively integrates complementary appearance, geometric, and spectral information, while a consistency learning module explicitly encourages feature coherence for genuine samples and emphasizes cross-modal discrepancies associated with spoof attacks. The consistency module is used only during training and removed during inference, avoiding additional deployment overhead.
Extensive experiments are conducted on CASIA-SURF, WMCA, CelebA-Spoof, and MSU-MFSD using intra-dataset, cross-dataset, ablation, feature-space, modality-consistency, and statistical analyses. The proposed CMCL achieves an ACER of 0.70% and accuracy of 99.3% on CASIA-SURF, while achieving an ACER of 2.55% and accuracy of 97.9% on WMCA. The results demonstrate improved spoof detection performance, representation discrimination, and cross-dataset generalization compared with the evaluated state-of-the-art methods. The proposed framework provides a consistency-driven and computationally practical approach for robust face anti-spoofing.

Cite This Paper

Mudunuru Suneel, Kaja Krishna Mohan, Banothu Yedukondala Venkata Naga Raja Swamy, Seva Sreedhar Babu, Venkata Raghavendra Miriampally, P. Rama Koteswara Rao, Kama Ramudu, "Cross-Modal Consistency Learning for Robust Face Anti-Spoofing Using RGB, RGB-Derived Depth and NIR Representations", International Journal of Wireless and Microwave Technologies(IJWMT), Vol.16, No.5, pp. 398-427, 2026. DOI:10.5815/ijwmt.2026.05.24

Reference

[1]Z. Ming, M. Visani, M. M. Luqman, and J.-C. Burie, "A survey on anti-spoofing methods for facial recognition with RGB cameras of generic consumer devices," J. Imaging, vol. 6, no. 12, Art. no. 139, Dec. 2020, doi: 10.3390/jimaging6120139.
[2]Z. Yu et al., "Deep learning for face anti-spoofing: A survey," IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 5, pp. 5609–5631, May 2023, doi: 10.1109/TPAMI.2022.3215850. 
[3]H. Li et al., "Face anti-spoofing via deep local binary patterns," in Computer Vision – ACCV 2016, Cham: Springer, 2017, pp. 543–558, doi: 10.1007/978-3-319-54181-5_36. 
[4]Y. Atoum, Y. Liu, A. Amin, and A. K. Jain, "Face anti-spoofing using patch and depth-based CNNs," in Proc. IEEE Int. Jt. Conf. Biometrics (IJCB), Denver, CO, USA, 2017, pp. 319–328, doi: 10.1109/BTAS.2017.8272713. 
[5]X. Sun, L. Huang, and C. Liu, "Multimodal face spoofing detection via RGB-D images," in Proc. 24th Int. Conf. Pattern Recognit. (ICPR), Beijing, China, 2018, pp. 2221-2226, doi: 10.1109/ICPR.2018.8545849.
[6]Y. Liu, A. Jourabloo, and X. Liu, "Learning deep models for face anti-spoofing: Binary or auxiliary supervision," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Salt Lake City, UT, USA, 2018, pp. 389–398, doi: 10.1109/CVPR.2018.00048. 
[7]Z. Yu, C. Zhao, Z. Wang, Y. Qin, Z. Su, X. Li, F. Zhou, and G. Zhao, "Searching central difference convolutional networks for face anti-spoofing," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Seattle, WA, USA, 2020, pp. 5294-5304, doi: 10.1109/CVPR42600.2020.00534.
[8]F. Jiang, P. Liu, and X. Zhou, "Multilevel fusing paired visible light and near-infrared spectral images for face anti-spoofing," Pattern Recognit. Lett., vol. 128, pp. 30-37, Dec. 2019, doi: 10.1016/j.patrec.2019.08.008.
[9]R. Shao, X. Lan, J. Li, and P. C. Yuen, "Multi-adversarial domain generalization for face anti-spoofing," in Proc. AAAI Conf. Artif. Intell., vol. 33, no. 01, 2019, pp. 8881–8888, doi: 10.1609/aaai.v33i01.33016843. 
[10]H. Nam, H. Lee, J. Park, W. Yoon, and D. Yoo, "Reducing domain gap by reducing style bias," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2021, pp. 8686-8695, doi: 10.1109/CVPR46437.2021.00858.
[11]Z. Yu et al., "Searching central difference convolutional networks for face anti-spoofing," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Seattle, WA, USA, 2020, pp. 5294–5304, doi: 10.1109/CVPR42600.2020.00534. 
[12]S. Zhang et al., "CASIA-SURF: A large-scale multi-modal benchmark for face anti-spoofing," IEEE Trans. Biometrics, Behav., Identity Sci., vol. 2, no. 2, pp. 182–193, Apr. 2020, doi: 10.1109/TBIOM.2020.2973001. 
[13]A. Liu et al., "Cross-ethnicity face anti-spoofing recognition challenge: A review," IET Biometrics, vol. 10, no. 1, pp. 24–43, 2021, doi: 10.1049/bme2.12002. 
[14]A. George et al., "Biometric face presentation attack detection with multi-channel convolutional neural network," IEEE Trans. Inf. Forensics Security, vol. 15, pp. 42–55, 2020, doi: 10.1109/TIFS.2019.2916652. 
[15]Y. Liu, J. Stehouwer, A. Jourabloo, and X. Liu, "Deep tree learning for zero-shot face anti-spoofing," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Long Beach, CA, USA, 2019, pp. 4675–4684, doi: 10.1109/CVPR.2019.00481. 
[16]M. Tan and Q. V. Le, "EfficientNet: Rethinking model scaling for convolutional neural networks," in Proc. 36th Int. Conf. Mach. Learn. (ICML), 2019, pp. 6105–6114, doi: 10.48550/arXiv.1905.11946. 
[17]A. Howard et al., "Searching for MobileNetV3," in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Seoul, South Korea, 2019, pp. 1314–1324, doi: 10.1109/ICCV.2019.00140. 
[18]A. Dosovitskiy et al., "An image is worth 16x16 words: Transformers for image recognition at scale," in Proc. Int. Conf. Learn. Represent. (ICLR), 2021, doi: 10.48550/arXiv.2010.11929. 
[19]A. Vaswani et al., "Attention is all you need," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 30, 2017, pp. 5998–6008, doi: 10.48550/arXiv.1706.03762. 
[20]I. Goodfellow et al., "Generative adversarial nets," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 27, 2014, pp. 2672–2680, doi: 10.48550/arXiv.1406.2661. 
[21]K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Las Vegas, NV, USA, 2016, pp. 770–778, doi: 10.1109/CVPR.2016.90. 
[22]J. Deng, J. Guo, N. Xue, and S. Zafeiriou, "ArcFace: Additive angular margin loss for deep face recognition," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Long Beach, CA, USA, 2019, pp. 4685–4694, doi: 10.1109/CVPR.2019.00482. 
[23]Y. Sun, C. Cheng, Y. Zhang, C. Zhang, L. Zheng, Z. Wang, and Y. Wei, "Circle loss: A unified perspective of pair similarity optimization," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Seattle, WA, USA, 2020, pp. 6397–6406, doi: 10.1109/CVPR42600.2020.00643. 
[24]M. Long, Y. Cao, J. Wang, and M. I. Jordan, "Learning transferable features with deep adaptation networks," in Proc. 32nd Int. Conf. Mach. Learn. (ICML), 2015, pp. 97–105, doi: 10.48550/arXiv.1502.02791. 
[25]Y. Ganin and V. Lempitsky, "Unsupervised domain adaptation by backpropagation," in Proc. 32nd Int. Conf. Mach. Learn. (ICML), 2015, pp. 1180–1189, doi: 10.48550/arXiv.1409.7495. 
[26]K. Zhou, Y. Yang, Y. Hospedales, and T. Xiang, "Domain generalization with MixStyle," in Proc. Int. Conf. Learn. Represent. (ICLR), 2021, doi: 10.48550/arXiv.2104.02008. 
[27]E. Tzeng, J. Hoffman, K. Saenko, and T. Darrell, "Adversarial discriminative domain adaptation," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Honolulu, HI, USA, 2017, pp. 2962–2971, doi: 10.1109/CVPR.2017.316. 
[28]J. P. Appadurai, R. S. N. Noella, M. S. Jacob, K. Arunasakthi, B. S. K. Devi, and M. K. Singh, "An effective biometric medical image watermarking system designed for e-Health application," in Machine Learning Algorithms for Data Security and Healthcare Applications, Cham: Springer, 2025, pp. 310–317, doi: 10.1007/978-3-031-75861-4_27. 
[29]T. Shen, Y. Huang, and Z. Tong, "FaceBagNet: Bag-of-local-features model for multi-modal face anti-spoofing," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), Long Beach, CA, USA, 2019, pp. 1611-1616, doi: 10.1109/CVPRW.2019.00203.
[30]K. Patel, H. Han, and A. K. Jain, "Secure face unlock: Spoof detection on smartphones," IEEE Trans. Inf. Forensics Security, vol. 11, no. 10, pp. 2268–2283, Oct. 2016, doi: 10.1109/TIFS.2016.2578288. 
[31]Y. Jia, J. Zhang, S. Shan, and X. Chen, "Unified unsupervised and semi-supervised domain adaptation network for cross-scenario face anti-spoofing," Pattern Recognit., vol. 115, Art. no. 107888, Jul. 2021, doi: 10.1016/j.patcog.2021.107888.
[32]H. Li, P. He, S. Wang, A. Rocha, X. Jiang, and A. C. Kot, "Learning generalized deep feature representation for face anti-spoofing," IEEE Trans. Inf. Forensics Security, vol. 13, no. 10, pp. 2639-2652, Oct. 2018, doi: 10.1109/TIFS.2018.2825949.
[33]H. Luo et al., "AlignedReID: Surpassing human-level performance in person re-identification," arXiv preprint arXiv:1711.08184, 2017, doi: 10.48550/arXiv.1711.08184. 
[34]X. Chen, S. Xu, Q. Ji, and S. Cao, "A dataset and benchmark towards multi-modal face anti-spoofing under surveillance scenarios," IEEE Access, vol. 9, pp. 28140-28155, 2021, doi: 10.1109/ACCESS.2021.3052728.
[35]Z. Boulkenafet, J. Komulainen, and A. Hadid, "Face anti-spoofing based on color texture analysis," in Proc. IEEE Int. Conf. Image Process. (ICIP), Quebec City, QC, Canada, 2015, pp. 2636–2640, doi: 10.1109/ICIP.2015.7351280. 
[36]A. Liu, Z. Tan, J. Wan, S. Escalera, G. Guo, and S. Z. Li, "CASIA-SURF CeFA: A benchmark for multi-modal cross-ethnicity face anti-spoofing," in Proc. IEEE Winter Conf. Appl. Comput. Vis. (WACV), Waikoloa, HI, USA, 2021, pp. 1178-1186, doi: 10.1109/WACV48630.2021.00122.
[37]A. Liu et al., "Static and dynamic fusion for multi-modal cross-ethnicity face anti-spoofing," arXiv preprint arXiv:1912.02340, 2019, doi: 10.48550/arXiv.1912.02340. 
[38]A. Liu et al., "FM-ViT: Flexible modal vision transformers for face anti-spoofing," IEEE Trans. Inf. Forensics Security, vol. 18, pp. 4775–4786, 2023, doi: 10.1109/TIFS.2023.3296330. 
[39]Y. Zhang et al., "CelebA-Spoof: Large-scale face anti-spoofing dataset with rich annotations," in Computer Vision – ECCV 2020, Cham: Springer, 2020, pp. 70–85, doi: 10.1007/978-3-030-58610-2_5. 
[40]J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, "ImageNet: A large-scale hierarchical image database," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Miami, FL, USA, 2009, pp. 248–255, doi: 10.1109/CVPR.2009.5206848. 
[41]D. Wen, H. Han, and A. K. Jain, "Face spoof detection with image distortion analysis," IEEE Trans. Inf. Forensics Security, vol. 10, no. 4, pp. 746–761, Apr. 2015, doi: 10.1109/TIFS.2015.2400395. 
[42]A. Mikołajczyk and M. Grochowski, "Data augmentation for improving deep learning in image classification problem," in Proc. Int. Interdisciplinary PhD Workshop (IIPhDW), Swinoujscie, Poland, 2018, pp. 117–122, doi: 10.1109/IIPHDW.2018.8388338. 
[43]I. Alhashim and P. Wonka, "High quality monocular depth estimation via transfer learning," arXiv preprint arXiv:1812.11941, 2018, doi: 10.48550/arXiv.1812.11941. 
[44]G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, "Densely connected convolutional networks," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Honolulu, HI, USA, 2017, pp. 4700–4708, doi: 10.1109/CVPR.2017.243.