IJEM Vol. 16, No. 5, 8 Oct. 2026
Cover page and Table of Contents: PDF (size: 1019KB)
PDF (1019KB), PP.312-329
Views: 0 Downloads: 0
Infrared and visible image fusion, Night-time intrusion detection, EfficientDet D3, Bidirectional feature pyramid network, Multimodal surveillance, Real-time object detection, Attention-guided fusion
Night-time perimeter monitoring in low illumination remains challenging because cluttered terrain, partial occlusion, and thermal crossover conditions simultaneously distort object boundaries and visual cues. Visible spectrum sensing loses discriminative texture at night, while infrared sensing preserves thermal salience but lacks structural context. Conventional pipelines that rely on single-modality detection or simple blending commonly produce unstable recall or avoidable false alarms when the background changes. The present paper presents an EfficientDet Fusion Intrusion Detector (EFID) which combines registered infrared and visible frames using an attention-guided weighting module and a spatially adaptive activity-weighted pixel fusion step. The feature-level attention weights estimate the reliability of each modality and guide the activity-weighted pixel fusion stage. The fused RGB image is resized from 1024 × 768 to 896 × 896 before being processed by EfficientDet-D3. The resultant fusion retains visible structural edges, while infrared target evidence is simultaneously enhanced. The resulting three-channel representation is processed by an EfficientDet D3 detector augmented with a bidirectional feature pyramid network (BiFPN) for multi-scale localisation under night conditions and clutter. Experiments use the public Multi-scenario Multi-modality Fusion and Detection (M³FD) benchmark containing 4,200 aligned infrared and visible pairs at 1024×768 resolution, annotated with 33,603 bounding boxes across six classes. The proposed EFID achieves 91.7% mAP@0.5 and 67.3% mAP@0.5:0.95, with 93.4% precision and 89.8% recall at 34.2 FPS on an NVIDIA RTX 3080. Systematic ablation confirms the independent contribution of the attention module, the activity-weighted pixel fusion term, and each BiFPN iteration. Cross-dataset experiments demonstrate that zero-shot transfer retains 84.1% mAP@0.5, rising to 86.0% with 10% target fine-tuning. The results indicate practical night robustness and real-time feasibility, while extreme occlusion and calibration drift remain limiting factors that motivate alignment-aware training and lightweight deployment optimisation as future work.
Rithick S., Jenefa A., Abirami M. K., Sheeba Merlin, Antony Taurshia, Lincy A., "Spatially Adaptive Infrared–visible Fusion with EfficientDet for Night-time Multi-Class Intrusion Detection", International Journal of Engineering and Manufacturing (IJEM), Vol.16, No.5, pp. 312-329, 2026. DOI:10.5815/ijem.2026.05.17
[1]Liu, Jinyuan, Xin Fan, Zhanbo Huang, Guanyao Wu, Risheng Liu, Wei Zhong, and Zhongxuan Luo. "Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection." In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 5792-5801. 2022.
[2]Tan, Mingxing, Ruoming Pang, and Quoc V. Le. "Efficientdet: Scalable and efficient object detection." In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10778-10787. 2020.
[3]Tang, Linfeng, Jiteng Yuan, Hao Zhang, Xingyu Jiang, and Jiayi Ma. "PIAFusion: A progressive infrared and visible image fusion network based on illumination aware." Information Fusion 83 (2022): 79-92.doi: 10.1016/j.inffus.2022.03.007.
[4]Roy, Swalpa Kumar, et al. "Multimodal fusion transformer for remote sensing image classification." IEEE Transactions on Geoscience and Remote Sensing 61 (2023): 1-20.doi: 10.1109/TGRS.2023.3286826.
[5]Ma, Rui, Yi Hou, Chenxuan Li, Huizhu Jia, and Xiaodong Xie. "Scene-adaptive unsupervised crowd counting for video surveillance." IEEE Transactions on Circuits and Systems for Video Technology (2025).doi: 10.1109/TCSVT.2025.3540850
[6]Ni, Peizhou, Weiming Hu, Jinchao Hu, Xiong Chen, and Yan Liu. "Illumination robust semantic segmentation based on cross-dimensional multispectral edge fusion in dynamic traffic scenes." IEEE Access 12 (2024): 171589-171600.doi: 10.1109/ACCESS.2024.3498896.
[7]Santhiya, P., Immanuel JohnRaja Jebadurai, Getzi Jeba Leelipushpam Paulraj, and E. Donald Jawahar. "Implementing ViT Models for Traffic Sign Detection in Autonomous Driving Systems." In 2024 5th International Conference on Recent Trends in Computer Science and Technology (ICRTCST), pp. 382-387. IEEE, 2024.
[8]Yang, Sen, Ze Feng, Zhicheng Wang, Yanjie Li, Shoukui Zhang, Zhibin Quan, Shu-tao Xia, and Wankou Yang. "Detecting and grouping keypoints for multi-person pose estimation using instance-aware attention." Pattern recognition 136 (2023): 109232.doi: 10.1016/j.patcog.2022.109232.
[9]Yang, Guang, Jie Li, Hanxiao Lei, and Xinbo Gao. "A multi-scale information integration framework for infrared and visible image fusion." Neurocomputing 600 (2024): 128116.doi: 10.1016/j.neucom.2024.128116.
[10]Taurshia, A., Kuriakose, B. M., Vijaya Kumar, E. N., & Lincy, A. (2025). Advancing Road Maintenance with EfficientDet-based Pothole Monitoring. Serbian Journal of Electrical Engineering, 22(1), 57.doi: 10.2298/SJEE2501057J
[11]Li, Hui, and Xiao-Jun Wu. "DenseFuse: A fusion approach to infrared and visible images." IEEE transactions on image processing 28.5 (2019): 2614-2623.doi: 10.1109/TIP.2018.2887342.
[12]Li, Hui, Xiao-Jun Wu, and Josef Kittler. "RFN-Nest: An end-to-end residual fusion network for infrared and visible images." Information Fusion 73 (2021): 72-86.doi: 10.1016/j.inffus.2021.02.023.
[13]Ma, Jiayi, Han Xu, Junjun Jiang, Xiaoguang Mei, and Xiao-Ping Zhang. "DDcGAN: A dual-discriminator conditional generative adversarial network for multi-resolution image fusion." IEEE Transactions on Image Processing 29 (2020): 4980-4995.doi: 10.1109/TIP.2020.2977573.
[14]Xu, Han, Jiayi Ma, Zhuliang Le, Junjun Jiang, and Xiaojie Guo. "Fusiondn: A unified densely connected network for image fusion." In Proceedings of the AAAI conference on artificial intelligence, vol. 34, no. 07, pp. 12484-12491. 2020.
[15]Santhiya, P., Jebadurai, I. J., Paulraj, G. J. L., & Karan, S. K. (2024, June). Deep Vision: Lane Detection in ITS: A Deep Learning Segmentation Perspective. In 2024 Second International Conference on Inventive Computing and Informatics (ICICI) (pp. 21-26). IEEE.
[16]Xiang, Xiantai, Guangyao Zhou, Ben Niu, Zongxu Pan, Lijia Huang, Wenshuai Li, Zixiao Wen, Jiamin Qi, and Wanxin Gao. "Infrared-visible image fusion meets object detection: Towards unified optimization for multimodal perception." Remote Sensing 17, no. 21 (2025): 3637.doi: 10.3390/rs17213637.
[17]Gallagher, James E., and Edward J. Oughton. "Surveying you only look once (YOLO) multispectral object detection advancements, applications, and challenges." IEEE Access 13 (2025): 7366-7395.doi: 10.1109/ACCESS.2025.3526458.
[18]Wang, Qinghua, et al. "DICFusion: Infrared and Visible Image Fusion via a Deep Integrated and Semantic-Coordinated Network." IEEE Transactions on Circuits and Systems for Video Technology (2025).doi: 10.1109/TCSVT.2025.3625990
[19]Madhasu, Nithya, and Sagar Dhanraj Pande. "Revolutionizing wildlife protection: A novel approach combining deep learning and night-time surveillance." Multimedia Tools and Applications 84.21 (2025): 23465-23499.doi: 10.1007/s11042-024-19876-4.
[20]Joshua Samuel, Giri Balan, and Joshua Premkumar. "Enhancing public safety through license plate recognition for counter terrorism through deep learning technique." In 2023 4th International Conference on Signal Processing and Communication (ICSPC), pp. 96-100. IEEE, 2023.
[21]Tu, Haiyan, Xiaoyue Tan, Zhongping Yin, Xuegang Zhang, Kang Yang, Zhengkun Qiu, and Xiujuan Zheng. "Non-intrusive fatigue detection based on multidomain features of sitting pressure and machine learning." IEEE Sensors Journal (2025).doi: 10.1109/JSEN.2025.3544299.
[22]Vasantharaj, A., J. Immanuel Jebaraj, K. Kavinraj, and S. Sebin. "An Intelligent Deep Learning Approach for Wild Animal Detection and Alert System Using YOLOv11." In 2025 10th International Conference on Communication and Electronics Systems (ICCES), pp. 1519-1526. IEEE, 2025.
[23]Wang, Jiali, Zhengyu Xie, Yong Qin, Xiaoqiang Zhang, Zengqing Wang, and Li Wang. "Boundary Dynamic Perception Detection Network for Low Light Railway Environment." IEEE Sensors Journal (2025).doi: 10.1109/JSEN.2025.3543860.
[24]Rocha, Pedro Daniel, Fernando JP Lopes, and Luís A. Da Silva Cruz. "Automating Electrical Grid Asset Inspection: From Current Challenges to Future Directions." IEEE Access 13 (2025): 201392-201438.doi: 10.1109/ACCESS.2025.3636718.
[25]Aman, Serge Stéphane, Behou Gérard N’guessan, Kouadio Prosper Kimou, and Tiemoman Kone. "AHOD: adaptive hybrid object detector for context-aware and real-time object detection in complex environments." Discover Applied Sciences 7, no. 11 (2025): 1254.doi: 10.1007/s42452-025-07784-7
[26]Ande, A., Mounikuttan, T., Anuj, M. D., Makros, G. J., Rejoice, G. R., & Shalini, T. M. (2023, March). Real-time rail safety: A deep convolutional neural network approach for obstacle detection on tracks. In 2023 4th International Conference on Signal Processing and Communication (ICSPC) (pp. 101-105). IEEE.
[27]Rampriya, R. S., Sahaya Beni Prathiba, L. Jerart Julus, Arikumar K. Selvaraj, and Joel JPC Rodrigues. "Fault detection and semantic segmentation on railway track using deep fusion model." IEEE Access 12 (2024): 136183-136201.doi: 10.1109/ACCESS.2024.3462822.
[28]Santhiya, P., Immanuel Johnraja Jebadurai and Getzi Jeba Leelipushpam Paulraj, "BIPOOLNET: An advanced UNet architecture for enhanced lane detection in autonomous vehicles." Intell. Decis. Technol. 18, no. 2 (2024): 743-757.doi: 10.3233/IDT-240162.
[29]Isaac J, Santhiya P, Archpaul J. HearTheSigns: AI-Based Road Sign Detection with Real-Time Voice Feedback for Blind. In2025 International Conference on Computing Technologies (ICOCT) 2025 Jun 13 (pp. 1-6). IEEE.
[30]RS, R., Al-Shehari, T., Nathan, S., A, J., R, S., P, S.P., Alfakih, T. and Alsalman, H., 2024. An unmanned aerial vehicle captured dataset for railroad segmentation and obstacle detection. Scientific Data, 11(1), p.1315.doi: 10.1038/s41597-024-03952-3.
[31]J. Hu, L. Shen, G. Sun, "Squeeze-and-excitation networks," in Proc. IEEE CVPR, 2018, pp. 7132-7141.
[32]T.-Y. Lin, P. Goyal, R. Girshick, K. He, P. Dollár, "Focal loss for dense object detection," in Proc. IEEE ICCV, 2017, pp. 2980-2988.
[33]Z. Zheng, P. Wang, W. Liu, J. Li, R. Ye, D. Ren, "Distance-IoU loss: Faster and better learning for bounding box regression," in Proc. AAAI, vol. 34, no. 7, pp. 12993-13000, 2020.