Work place: K J Somaiya School of Engineering (formerly known as K J Somaiya College of Engineering), Somaiya Vidyavihar University, Mumbai, India
E-mail: hetvi.shah6@somaiya.edu
Website: https://orcid.org/0009-0003-5562-3053
Research Interests:
Biography
Hetvi Rajesh Shah received her bachelor’s degree in Information Technology along with honors in artificial intelligence from K. J. Somaiya College of Engineering in 2024. Her research areas include machine learning, artificial intelligence, NLP, and full-stack development.
By Dipti Pawade Sonali Patil Riddhi Arya Hetvi Shah Diya Bakhai Ankit Jha
DOI: https://doi.org/10.5815/ijieeb.2026.04.09, Pub. Date: 8 Aug. 2026
Remote sensing images are complex, which makes it difficult to interpret and generate semantically appropriate textual description. To get a semantically relevant description, it is important to identify complex objects and understand the contextual relationships between them. In such cases, deriving contextually accurate information while maintaining semantic coherence is challenging. Therefore, a specifically designed model architecture is required to generate semantically relevant descriptions. This paper discusses a deep learning-based approach to generate remote sensing image descriptions using an end-to-end encoder-decoder model with soft attention. The UC Merced (UCM) dataset is used for training, which includes multiple captions per image capturing various scene aspects. To further assess the robustness and generalizability of the proposed approach, its performance is additionally evaluated on more complex datasets such as RSCID and Sydney Captions. This study presents an end-to-end CNN–LSTM encoder–decoder framework enhanced with soft attention for semantic description generation from remote sensing imagery. The framework employs a VGG16 encoder to extract a 4096-dimensional visual feature vector, which is projected into a 256-dimensional representation and processed by a 256-unit LSTM decoder. The soft attention mechanism dynamically computes attention weights using the encoder features and decoder hidden state, enabling the model to emphasize relevant visual information during word generation. Multiple CNN encoders and learning rates are evaluated with LSTM decoders, both with and without attention, on the UCM, RSCID, and Sydney Caption datasets. At a learning rate of 0.0001, VGG16–LSTM with soft attention achieves BLEU-4 (B4) scores of 0.6636, 0.6636, and 0.5864 on the UCM, RSCID, and Sydney Caption datasets, respectively, compared with 0.1643, 0.1647, and 0.1745 for VGG16–LSTM without attention. The results demonstrate that soft attention substantially improves description generation by strengthening visual–linguistic alignment and enabling more contextually relevant and semantically coherent descriptions across datasets with varying scene complexity.
[...] Read more.Subscribe to receive issue release notifications and newsletters from MECS Press journals