IJIGSP Vol. 18, No. 4, 8 Aug. 2026
Cover page and Table of Contents: PDF (size: 2024KB)
PDF (2024KB), PP.25-47
Views: 0 Downloads: 0
Computer-Generated Image, Computer Graphics, Image Forensics, Fake Image Detection
The rapid development in deep learning-based generative softwares and image rendering tools has led to generation of massive photorealistic digital content – Fake images, fake videos, fake speech, etc. Such fake digital media may result in communication of misinformation, forgery of digital data, and losing trustworthiness in the information source. This poses a significant challenge to the field of digital forensics’ techniques, to which our present work attempts to make a contribution, by addressing the problem of differentiating AI-generated images from real photographs, using transfer learning and multi-branch fusion model. We propose a multi-branch model that integrates two pre-trained Vision Transformer models (DINO (self-distillation with no labels) and Contrastive Language – Image Pretraining (CLIP)) to extract complementary global features, along with a forensic and a hand-crafted feature branch, which extract low-level discriminating cues. These features are complimentary to each other and hence contribute in improving the robustness and performance of the model. These features from the four branches are adaptively weighted and combined by a cross-attention module, to give a fused and rich embedding. The model is further optimized by using augmentation-invariant loss, center loss and supervised contrastive loss in addition to the cross-entropy loss function. This framework achieves improved accuracy of 95.83% on PRCG dataset, 96.79% on CIFAKE dataset and 99.15% on GenImage dataset as compared to baselines. It also achieved stable cross-generator performance and enhanced robustness against real world corruptions like Blur, Noise, Compression, and others. The experimental results show a good separability between the classes, and enhanced performance on publicly available datasets.
Venkata Satya Renuka Devi Bhamidipati, Srinivasa Rao Chanamallu, Sudheer Gopinathan, "A Multi-Branch Transformer Based Cross-Attention Framework for Computer-Generated Image Detection", International Journal of Image, Graphics and Signal Processing(IJIGSP), Vol.18, No.4, pp. 25-47, 2026. DOI:10.5815/ijigsp.2026.04.02
[1]Brock A, Donahue J, Simonyan K. Large Scale Gan Training For High Fidelity Natural Image Synthesis. Int. Conf. Learn. Represent., 2019. https://doi.org/10.48550/arXiv.1809.11096.
[2]Lu Z, Huang D, Bai L, Qu J, Wu C, Liu X, et al. Seeing is not always believing: Benchmarking Human and Model Perception of AI-Generated Images. Proc. 37th Int. Conf. Neural Inf. Process. Syst., New Orleans, LA, USA: Curran Associates Inc.; 2023, p. 25435–47. https://doi.org/10.48550/arXiv.2304.13023.
[3]Gangu Rama Naidu, Chanamallu Srinivasa Rao. A CNN based Discrimination between Natural and Computer Generated Images. Pmj 2024;35:580–9. https://doi.org/10.52783/pmj.v35.i2s.2948.
[4]Nataraj L, Mohammed TM, Manjunath BS, Chandrasekaran S, Flenner A, Bappy JH, et al. Detecting GAN generated Fake Images using Co-occurrence Matrices. Electron Imaging 2019;31:532-1-532–7. https://doi.org/10.2352/ISSN.2470-1173.2019.5.MWSF-532.
[5]Baskar C, Govindasamy GP, Anbalagan S, Roomi SMM. Computer Graphic and Photographic Image Classification Using Transfer Learning Approach. Trait Signal 2022;39:1267–73. https://doi.org/10.18280/ts.390419.
[6]Castillo Camacho I, Wang K. Convolutional neural network initialization approaches for image manipulation detection. Digit Signal Process 2022;122:103376. https://doi.org/10.1016/j.dsp.2021.103376.
[7]Wang K. Self-Supervised Learning for the Distinction between Computer-Graphics Images and Natural Images. Appl Sci 2023;13:1887. https://doi.org/10.3390/app13031887.
[8]Ng T-T, Chang S-F, Hsu J, Xie L, Tsui M-P. Physics-motivated features for distinguishing photographic images and computer graphics. Proc. 13th Annu. ACM Int. Conf. Multimed., Hilton Singapore: ACM; 2005, p. 239–48. https://doi.org/10.1145/1101149.1101192.
[9]Zhang R-S, Quan W-Z, Fan L-B, Hu L-M, Yan D-M. Distinguishing Computer-Generated Images from Natural Images Using Channel and Pixel Correlation. J Comput Sci Technol 2020;35:592–602. https://doi.org/10.1007/s11390-020-0216-9.
[10]Gangan MP, K A, L LV. Distinguishing Natural and Computer-Generated Images using Multi-Colorspace fused EfficientNet. J Inf Secur Appl 2022;68:103261. https://doi.org/10.1016/j.jisa.2022.103261.
[11]Quan W, Wang K, Yan D-M, Zhang X, Pellerin D. Learn with diversity and from harder samples: Improving the generalization of CNN-Based detection of computer-generated images. Forensic Sci Int Digit Investig 2020;35:301023. https://doi.org/10.1016/j.fsidi.2020.301023.
[12]Grommelt P, Weiss L, Pfreundt F-J, Keuper J. Fake or JPEG? Revealing Common Biases in Generated Image Detection Datasets. In: Del Bue A, Canton C, Pont-Tuset J, Tommasi T, editors. Comput. Vis. – ECCV 2024 Workshop, vol. 15644, Cham: Springer Nature Switzerland; 2025, p. 80–95. https://doi.org/10.1007/978-3-031-92089-9_6.
[13]Guillaro F, Cozzolino D, Sud A, Dufour N, Verdoliva L. TruFor: Leveraging All-Round Clues for Trustworthy Image Forgery Detection and Localization. 2023 IEEECVF Conf. Comput. Vis. Pattern Recognit. CVPR, Vancouver, BC, Canada: IEEE; 2023, p. 20606–15. https://doi.org/10.1109/CVPR52729.2023.01974.
[14]Wang B, Wang F, Wang J, Yan H, Zhang F, Zhou S, et al. Unifaked: Universal Fake Image Detection Via Multiple Ensemble Learning 2025. https://doi.org/10.2139/ssrn.5139156.
[15]Jeong Y, Kim D, Ro Y, Choi J. FrePGAN: Robust Deepfake Detection Using Frequency-Level Perturbations. Proc AAAI Conf Artif Intell 2022;36:1060–8. https://doi.org/10.1609/aaai.v36i1.19990.
[16]Tan C, Zhao Y, Wei S, Gu G, Wei Y. Learning on Gradients: Generalized Artifacts Representation for GAN-Generated Images Detection. 2023 IEEECVF Conf. Comput. Vis. Pattern Recognit. CVPR, Vancouver, BC, Canada: IEEE; 2023, p. 12105–14. https://doi.org/10.1109/CVPR52729.2023.01165.
[17]Veselska O, Ziubina R. Reversible image steganography using transformer-based latent embedding. Adv Sci Technol Res J 2025;19:148–64. https://doi.org/10.12913/22998624/204419.
[18]Chen Y, Mai T, Wang S, Kang X. Computer-Generated Image Forensics Based on Vision Transformer with High-Frequency Feature Enhancement. 2025 33rd Eur. Signal Process. Conf. EUSIPCO, Palermo, Italy: IEEE; 2025, p. 1218–22. https://doi.org/10.23919/EUSIPCO63237.2025.11226696.
[19]Lamichhane D. Advanced Detection of AI-Generated Images Through Vision Transformers. IEEE Access 2025;13:3644–52. https://doi.org/10.1109/ACCESS.2024.3522759.
[20]Caron M, Touvron H, Misra I, Jegou H, Mairal J, Bojanowski P, et al. Emerging Properties in Self-Supervised Vision Transformers. 2021 IEEECVF Int. Conf. Comput. Vis. ICCV, Montreal, QC, Canada: IEEE; 2021, p. 9630–40. https://doi.org/10.1109/ICCV48922.2021.00951.
[21]Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, et al. Learning Transferable Visual Models From Natural Language Supervision. Proc. 38th Int. Conf. Mach. Learn., vol. 139, PMLR; 2021, p. 8748--8763. https://doi.org/10.48550/arXiv.2103.00020.
[22]Tan C, Zhao Y, Wei S, Gu G, Liu P, Wei Y. Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain Learning. Proc AAAI Conf Artif Intell 2024;38:5052–60. https://doi.org/10.1609/aaai.v38i5.28310.
[23]Alam I, Muneer MS, Woo SS. UGAD: Universal Generative AI Detector utilizing Frequency Fingerprints. Proc. 33rd ACM Int. Conf. Inf. Knowl. Manag., 2024, p. 4332–40. https://doi.org/10.1145/3627673.3680085.
[24]Cozzolino D, Poggi G, Verdoliva L. Recasting Residual-based Local Descriptors as Convolutional Neural Networks: an Application to Image Forgery Detection. Proc. 5th ACM Workshop Inf. Hiding Multimed. Secur., Philadelphia Pennsylvania USA: ACM; 2017, p. 159–64. https://doi.org/10.1145/3082031.3083247.
[25]Ming-Kuei Hu. Visual pattern recognition by moment invariants. IEEE Trans Inf Theory 1962;8:179–87. https://doi.org/10.1109/TIT.1962.1057692.
[26]Zhu M, Chen H, Yan Q, Huang X, Lin G, Li W, et al. GenImage: A Million-Scale Benchmark for Detecting AI-Generated Image. Proc. 37th Int. Conf. Neural Inf. Process. Syst., New Orleans, LA, USA: 2023, p. 77771–82. https://doi.org/10.48550/arXiv.2306.08571.
[27]Bird JJ, Lotfi A. CIFAKE: Image Classification and Explainable Identification of AI-Generated Synthetic Images. IEEE Access 2024;12:15642–50. https://doi.org/10.1109/ACCESS.2024.3356122.
[28]Ng T-T, Chang S-F, Hsu J, Pepeljugoski M. Columbia Photographic Images and Photorealistic Computer Graphics Dataset. ADVENT, Columbia University; 2005.
[29]Nichol A, Dhariwal P, Ramesh A, Shyam P, Mishkin P, McGrew B, et al. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. Int. Conf. Mach. Learn., arXiv; 2022. https://doi.org/10.48550/arXiv.2112.10741.
[30]Gu S, Chen D, Bao J, Wen F, Zhang B, Chen D, et al. Vector Quantized Diffusion Model for Text-to-Image Synthesis. 2022 IEEECVF Conf. Comput. Vis. Pattern Recognit. CVPR, New Orleans, LA, USA: IEEE; 2022, p. 10686–96. https://doi.org/10.1109/CVPR52688.2022.01043.
[31]Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B. High-Resolution Image Synthesis with Latent Diffusion Models. 2022 IEEECVF Conf. Comput. Vis. Pattern Recognit. CVPR, New Orleans, LA, USA: IEEE; 2022, p. 10674–85. https://doi.org/10.1109/CVPR52688.2022.01042.
[32]Dhariwal P, Nichol A. Diffusion Models Beat GANs on Image Synthesis. NIPS’21, vol. 672, Red Hook, NY, USA: Curran Associates Inc.; 2021, p. 8780–94. https://doi.org/10.48550/arXiv.2105.05233.
[33]Islam MT, Lee IH, Alzahrani AI, Muhammad K. MEXFIC: A meta ensemble eXplainable approach for AI-synthesized fake image classification. Alex Eng J 2025;116:351–63. https://doi.org/10.1016/j.aej.2024.12.031.
[34]Bartos GE, Akyol S. Deep Learning for Image Authentication: A Comparative Study on Real and AI-Generated Image Classification. 18th Int. Symp. Appl. Inform. Relat. Areas, Szekesfehervar, Hungary: n.d.
[35]Zhu M, Chen H, Huang M, Li W, Hu H, Hu J, et al. GenDet: Towards Good Generalizations for AI-Generated Image Detection. ArXiv 2023;abs/2312.08880.
[36]Chen J, Yao J, Niu L. A Single Simple Patch is All You Need for AI-generated Image Detection 2024. https://doi.org/10.48550/arXiv.2402.01123.