IJIEEB Vol. 18, No. 4, 8 Aug. 2026
Cover page and Table of Contents: PDF (size: 1871KB)
PDF (1871KB), PP.135-151
Views: 0 Downloads: 0
Remote Sensing Image Description, CNN, LSTM, VGG16, Encoder-Decoder, Attention Mechanism
Remote sensing images are complex, which makes it difficult to interpret and generate semantically appropriate textual description. To get a semantically relevant description, it is important to identify complex objects and understand the contextual relationships between them. In such cases, deriving contextually accurate information while maintaining semantic coherence is challenging. Therefore, a specifically designed model architecture is required to generate semantically relevant descriptions. This paper discusses a deep learning-based approach to generate remote sensing image descriptions using an end-to-end encoder-decoder model with soft attention. The UC Merced (UCM) dataset is used for training, which includes multiple captions per image capturing various scene aspects. To further assess the robustness and generalizability of the proposed approach, its performance is additionally evaluated on more complex datasets such as RSCID and Sydney Captions. This study presents an end-to-end CNN–LSTM encoder–decoder framework enhanced with soft attention for semantic description generation from remote sensing imagery. The framework employs a VGG16 encoder to extract a 4096-dimensional visual feature vector, which is projected into a 256-dimensional representation and processed by a 256-unit LSTM decoder. The soft attention mechanism dynamically computes attention weights using the encoder features and decoder hidden state, enabling the model to emphasize relevant visual information during word generation. Multiple CNN encoders and learning rates are evaluated with LSTM decoders, both with and without attention, on the UCM, RSCID, and Sydney Caption datasets. At a learning rate of 0.0001, VGG16–LSTM with soft attention achieves BLEU-4 (B4) scores of 0.6636, 0.6636, and 0.5864 on the UCM, RSCID, and Sydney Caption datasets, respectively, compared with 0.1643, 0.1647, and 0.1745 for VGG16–LSTM without attention. The results demonstrate that soft attention substantially improves description generation by strengthening visual–linguistic alignment and enabling more contextually relevant and semantically coherent descriptions across datasets with varying scene complexity.
Dipti Pawade, Sonali Patil, Riddhi Arya, Hetvi Shah, Diya Bakhai, Ankit Jha, "Soft Attention Enhanced CNN and LSTM Based Framework for Semantic Description Generation of Remote Sensing Imagery", International Journal of Information Engineering and Electronic Business(IJIEEB), Vol.18, No.4, pp. 135-151, 2026. DOI:10.5815/ijieeb.2026.04.09
[1]D. Pawade, A. Sakhapara, C. Shah, J. Wala, A. Tripathi, and B. Shah, “Text Caption Generation Based on Lip Movement of Speaker in Video Using Neural Network,” in Advances in Computing and Data Sciences, M. Singh, P. K. Gupta, V. Tyagi, J. Flusser, T. Ören, and R. Kashyap, Eds., Singapore: Springer Singapore, 2019, pp. 313–322.
[2]H. Li, H. Wang, Y. Zhang, L. Li, and P. Ren, “Underwater image captioning: Challenges, models, and datasets,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 220, pp. 440–453, Feb. 2025, doi: 10.1016/j.isprsjprs.2024.12.002.
[3]O. Arshi and P. Dadure, “A Comprehensive Review of Image Caption Generation,” Multimed. Tools Appl., vol. 84, no. 25, pp. 29419–29471, Jul. 2025, doi: 10.1007/s11042-024-20095-0.
[4]S. Zou, Y. Wei, Y. Xie, M. Lao, and X. Luan, “Remote Sensing Image Change Captioning: A Comprehensive Review,” Int. J. Multimed. Inf. Retr., vol. 14, no. 3, pp. 1–20, Jul. 2025, doi: 10.1007/s13735-025-00375-7.
[5]J. Castro Lopes, J. L. Oliveira, and R. P. Lopes, “A Comprehensive Review of Synthetic Image Generation Methods in Remote Sensing,” Jul. 08, 2025, Taylor and Francis Ltd. doi: 10.1080/01431161.2025.2527373.
[6]D. Pawade, H. Shah, R. Arya, and A. Jha, “Encoder Decoder Based Remote Sensing Image Description Generation,” in Proceedings of International Conference on Recent Trends in Computing, R. P. Mahapatra, S. Roy, and P. Parwekar, Eds., Singapore: Springer Nature Singapore, 2025, pp. 243–251.
[7]O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and Tell: A Neural Image Caption Generator,” Apr. 2015, [Online]. Available: http://arxiv.org/abs/1411.4555
[8]D. Sun, Y. Bao, J. Liu, and X. Cao, “A Lightweight Sparse Focus Transformer for Remote Sensing Image Change Captioning,” IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., vol. 17, pp. 18727–18738, 2024, doi: 10.1109/JSTARS.2024.3471625.
[9]J. D. Silva, J. Magalhães, D. Tuia, and B. Martins, “Large Language Models for Captioning and Retrieving Remote Sensing Images,” Feb. 2024, [Online]. Available: http://arxiv.org/abs/2402.06475
[10]S. Das, D. Mundra, P. Dayal, and R. Sharma, “A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning,” Jun. 2025, [Online]. Available: http://arxiv.org/abs/2506.09429
[11]Z. Shi and Z. Zou, “Can a Machine Generate Humanlike Language Descriptions for a Remote Sensing Image?,” IEEE Transactions on Geoscience and Remote Sensing, vol. 55, no. 6, pp. 3623–3634, 2017, doi: 10.1109/TGRS.2017.2677464.
[12]W. Huang, Q. Wang, and X. Li, “Denoising-Based Multiscale Feature Fusion for Remote Sensing Image Captioning,” IEEE Geoscience and Remote Sensing Letters, vol. 18, no. 3, pp. 436–440, 2021, doi: 10.1109/LGRS.2020.2980933.
[13]J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille, “Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN),” Jun. 2015, [Online]. Available: http://arxiv.org/abs/1412.6632
[14]G. Hoxha and F. Melgani, “Remote Sensing Image Captioning with SVM-Based Decoding,” in International Geoscience and Remote Sensing Symposium (IGARSS), Institute of Electrical and Electronics Engineers Inc., Sep. 2020, pp. 6734–6737. doi: 10.1109/IGARSS39084.2020.9323651.
[15]G. Sumbul, S. Nayak, and B. Demir, “SD-RSIC: Summarization Driven Deep Remote Sensing Image Captioning,” IEEE Transactions on Geoscience and Remote Sensing, vol. 59, no. 8, pp. 6922–6934, Aug. 2021.
[16]B. Wang, X. Lu, X. Zheng, and X. Li, “Semantic Descriptions of High-Resolution Remote Sensing Images,” IEEE Geoscience and Remote Sensing Letters, vol. 16, no. 8, pp. 1274–1278, Aug. 2019, doi: 10.1109/LGRS.2019.2893772.
[17]X. Lu, B. Wang, and X. Zheng, “Sound Active Attention Framework for Remote Sensing Image Captioning,” IEEE Transactions on Geoscience and Remote Sensing, vol. 58, no. 3, pp. 1985–2000, 2020, doi: 10.1109/TGRS.2019.2951636.
[18]S. Wu, X. Zhang, X. Wang, C. Li, and L. Jiao, “Scene Attention Mechanism for Remote Sensing Image Caption Generation,” in 2020 International Joint Conference on Neural Networks (IJCNN), 2020, pp. 1–7. doi: 10.1109/IJCNN48605.2020.9207381.
[19]H. Chen and Z. Shi, “A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,” Remote Sens. (Basel)., vol. 12, no. 10, May 2020, doi: 10.3390/rs12101662.
[20]U. Zia, M. Mohsin Riaz, and A. Ghafoor, “Transforming remote sensing images to textual descriptions,” International Journal of Applied Earth Observation and Geoinformation, vol. 108, p. 102741, 2022, doi: https://doi.org/10.1016/j.jag.2022.102741.
[21]Z. Zhang, W. Zhang, M. Yan, X. Gao, K. Fu, and X. Sun, “Global Visual Feature and Linguistic State Guided Attention for Remote Sensing Image Captioning,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, 2022, doi: 10.1109/TGRS.2021.3132095.
[22]Z. Chen, J. Wang, A. Ma, and Y. Zhong, “TypeFormer: Multiscale Transformer With Type Controller for Remote Sensing Image Caption,” IEEE Geoscience and Remote Sensing Letters, vol. 19, p. 3192062, Jan. 2022, doi: 10.1109/LGRS.2022.3192062.
[23]C. Liu, R. Zhao, H. Chen, Z. Zou, and Z. Shi, “Remote Sensing Image Change Captioning With Dual-Branch Transformers: A New Method and a Large Scale Dataset,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, 2022, doi: 10.1109/TGRS.2022.3218921.
[24]H. Kandala, S. Saha, B. Banerjee, and X. X. Zhu, “Exploring Transformer and Multilabel Classification for Remote Sensing Image Captioning,” IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1–5, 2022, doi: 10.1109/LGRS.2022.3198234.
[25]S. He, W. Liao, H. R. Tavakoli, M. Yang, B. Rosenhahn, and N. Pugeault, “Image Captioning through Image Transformer,” Oct. 2020, [Online]. Available: http://arxiv.org/abs/2004.14231
[26]Yi Yang and Shawn Newsam, “Bag-of-visual-words and spatial extensions for land-use classification,” in 18th ACM SIGSPATIAL International Symposium on Advances in Geographic Information Systems, San Jose, CA, USA: ACM Digital Library, Nov. 2013.
[27]C. S. NagaDurga and T. Anuradha, “Attention-Based Comparison of Automatic Image Caption Generation Encoders,” in Advances in Micro-Electronics, Embedded Systems and IoT, V. V. S. S. S. Chakravarthy, W. Flores-Fuentes, V. Bhateja, and B. N. Biswal, Eds., Singapore: Springer Nature Singapore, 2022, pp. 157–167.
[28]S. Degadwala, D. Vyas, H. Biswas, U. Chakraborty, and S. Saha, “Image Captioning Using Inception V3 Transfer Learning Model,” in 2021 6th International Conference on Communication and Electronics Systems (ICCES), 2021, pp. 1103–1108. doi: 10.1109/ICCES51350.2021.9489111.
[29]S. C. Kumar, M. Hemalatha, S. B. Narayan, and P. Nandhini, “Region Driven Remote Sensing Image Captioning,” Procedia Comput. Sci., vol. 165, pp. 32–40, 2019, doi: https://doi.org/10.1016/j.procs.2020.01.067.
[30]S. Mascarenhas and M. Agarwal, “A comparison between VGG16, VGG19 and ResNet50 architecture frameworks for Image Classification,” in 2021 International Conference on Disruptive Technologies for Multi-Disciplinary Research and Applications (CENTCON), 2021, pp. 96–99. doi: 10.1109/CENTCON52345.2021.9687944.
[31]K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556, 2014.
[32]F. A. Gers, J. Schmidhuber, and F. Cummins, “Learning to Forget: Continual Prediction with LSTM,” Neural Comput., vol. 12, no. 10, pp. 2451–2471, 2000, doi: 10.1162/089976600300015015.
[33]M. Sundermeyer, R. Schlüter, and H. Ney, “LSTM Neural Networks for Language Modeling,” in INTERSPEECH 2012 ISCA’s 13th Annual Conference Portland, OR, USA, USA, Oct. 2012, pp. 194–197.
[34]Y. Chu, X. Yue, L. Yu, M. Sergei, and Z. Wang, “Automatic Image Captioning Based on ResNet50 and LSTM with Soft Attention,” Wirel. Commun. Mob. Comput., vol. 2020, no. 1, p. 8909458, Jan. 2020, doi: https://doi.org/10.1155/2020/8909458.
[35]K. Cheng, J. Liu, R. Mao, Z. Wu, and E. Cambria, “CSA-RSIC: Cross-Modal Semantic Alignment for Remote Sensing Image Captioning,” IEEE Geoscience and Remote Sensing Letters, vol. 22, 2025, doi: 10.1109/LGRS.2025.3601114.
[36]Q. Wang, W. Huang, X. Zhang, and X. Li, “Word-Sentence Framework for Remote Sensing Image Captioning,” IEEE Transactions on Geoscience and Remote Sensing, vol. 59, no. 12, pp. 10532–10543, Dec. 2021, doi: 10.1109/TGRS.2020.3044054.