Dipti Pawade

Work place: K. J. Somaiya College of Engineering, Department of IT, Mumbai, Maharashtra, India

E-mail: diptipawade@somaiya.edu

Website:

Research Interests: Network Security, Application Security, Computational Learning Theory

Biography

Dipti Yogesh Pawade received B.E. degree in Computer Science and Engineering from Sant Gadge Baba University in 2009 and M.E. degree in Embedded System and Computing from G. H. Raisoni College of Engineering, Nagpur in 2012. Since 2012 she is an Assistant Professor in the Department of Information Technology at K J Somaiya College of Engineering, Vidyavihar, Mumbai. Her interest includes Machine learning, Web Security, and Web Application Development.

Author Articles
Soft Attention Enhanced CNN and LSTM Based Framework for Semantic Description Generation of Remote Sensing Imagery

By Dipti Pawade Sonali Patil Riddhi Arya Hetvi Shah Diya Bakhai Ankit Jha

DOI: https://doi.org/10.5815/ijieeb.2026.04.09, Pub. Date: 8 Aug. 2026

Remote sensing images are complex, which makes it difficult to interpret and generate semantically appropriate textual description. To get a semantically relevant description, it is important to identify complex objects and understand the contextual relationships between them. In such cases, deriving contextually accurate information while maintaining semantic coherence is challenging. Therefore, a specifically designed model architecture is required to generate semantically relevant descriptions. This paper discusses a deep learning-based approach to generate remote sensing image descriptions using an end-to-end encoder-decoder model with soft attention. The UC Merced (UCM) dataset is used for training, which includes multiple captions per image capturing various scene aspects. To further assess the robustness and generalizability of the proposed approach, its performance is additionally evaluated on more complex datasets such as RSCID and Sydney Captions. This study presents an end-to-end CNN–LSTM encoder–decoder framework enhanced with soft attention for semantic description generation from remote sensing imagery. The framework employs a VGG16 encoder to extract a 4096-dimensional visual feature vector, which is projected into a 256-dimensional representation and processed by a 256-unit LSTM decoder. The soft attention mechanism dynamically computes attention weights using the encoder features and decoder hidden state, enabling the model to emphasize relevant visual information during word generation. Multiple CNN encoders and learning rates are evaluated with LSTM decoders, both with and without attention, on the UCM, RSCID, and Sydney Caption datasets. At a learning rate of 0.0001, VGG16–LSTM with soft attention achieves BLEU-4 (B4) scores of 0.6636, 0.6636, and 0.5864 on the UCM, RSCID, and Sydney Caption datasets, respectively, compared with 0.1643, 0.1647, and 0.1745 for VGG16–LSTM without attention. The results demonstrate that soft attention substantially improves description generation by strengthening visual–linguistic alignment and enabling more contextually relevant and semantically coherent descriptions across datasets with varying scene complexity.

[...] Read more.
Closed Domain Question Answering System Tailored for Crime Events Using Deep Learning for Both Statistical and Contextualized Responses

By Dipti Pawade Sonali Patil Chaitanya Bandiwdekar Siddhesh Bagwe Pooja Kaulgud Aditi Kulkarni

DOI: https://doi.org/10.5815/ijieeb.2024.04.04, Pub. Date: 8 Aug. 2024

The goal of the question-answering system is to respond to user queries expressed in natural language. Unlike search engines, the closed domain question answering systems are specialized to specific domains, providing concise and precise answers often derived from structured data. This paper focuses on a question-answering system tailored for crime events, capable of addressing both statistical and contextual inquiries. In terms of crime statistics, the fine-tuned GPT-3 model outperforms the USE, TAPAS, TAPEX, and GPT-3 models, while for context-based crime-related queries, the fine-tuned RoBERTa model surpasses the BERT and RoBERTa models. This system is capable of providing the responses in natural language format, supplemented with relevant data visualizations. The models are train on Q2A and NewsQA datasets while it is tested on NCRB and NewsTimes datasets. The Q2A and NCRB datasets are used for statistical queries while NewsQA and NewsTimes datasets are used for contextual inquiries. The paper presents an analysis of various models and showcases results for sample case studies. Such a system can prove valuable in applications where users seek to study criminal cases or gather pertinent insights for specific cases. Furthermore, it can assist in understanding patterns and trends in criminal events, particularly concerning geospatial information. Linking crime event-based question-answering systems to geospatial information facilitates exploration of niche areas and furnishes precise details about local crime with minimal hype and hence worth exploring.

[...] Read more.
A Novel Approach for Video Inpainting Using Autoencoders

By Irfan Siddavatam Ashwini Dalvi Dipti Pawade Akshay Bhatt Jyeshtha Vartak Arnav Gupta

DOI: https://doi.org/10.5815/ijieeb.2021.06.05, Pub. Date: 8 Dec. 2021

Inpainting is a task undertaken to fill in damaged or missing parts of an image or video frame, with believable content. The aim of this operation is to realistically complete images or frames of videos for a variety of applications such as conservation and restoration of art, editing images and videos for aesthetic purposes, but might cause malpractices such as evidence tampering. From the image and video editing perspective, inpainting is used mainly in the context of generating content to fill the gaps left after removing a particular object from the image or the video. Video Inpainting, an extension of Image Inpainting, is a much more challenging task due to the constraint added by the time dimension. Several techniques do exist that achieve the task of removing an object from a given video, but they are still in a nascent stage. The major objective of this paper is to study the available approaches of inpainting and propose a solution to the limitations of existing inpainting techniques. After studying existing inpainting techniques, we realized that most of them make use of a ground truth frame to generate plausible results. A 'ground truth' frame is an image without the target object or in other words, an image that provides maximum information about the background, which is then used to fill spaces after object removal. In this paper, we propose an approach where there is no requirement of a 'ground truth' frame, provided that the video has enough contexts available about the background that is to be recreated. We would be using frames from the video in hand, to gather context for the background. As the position of the target object to be removed will vary from one frame to the next, each subsequent frame will reveal the region that was initially behind the object, and provide more information about the background as a whole. Later, we have also discussed the potential limitations of our approach and some workarounds for the same, while showing the direction for further research.

[...] Read more.
Blockchain Based Secure Traffic Police Assistant System

By Dipti Pawade Avani Sakhapara Raj shah Siby Thampi Vignesh Vaid

DOI: https://doi.org/10.5815/ijeme.2020.06.05, Pub. Date: 8 Dec. 2020

There are large numbers of vehicles in the populated country like India. It’s a very common scenario that traffic police came across some vehicle random vehicle and had some doubt in mind but do not have in hand information about that vehicle and end up leaving that thought. Sometimes this may result in some disaster. With the advent of technology, there are mobile applications and web based systems are available to ease up the process by which traffic police can fine the vehicle owner or people can pay the fine online. But yet there is no system is available through which traffic police can get all the details about the particular vehicle. This motivated us to design and developed an application thorough which traffic police can get all the information right from owner of the vehicle to its RC book and insurance status on just one click. Looking at the chances of data tampering, we have also played an attention to the data security and used blockchain for creating distributed, robust and tempered proof system. In this paper we have discussed traffic police assistance system, which can scan the vehicle number plate, identify the number and provide the all the information and documents stored against that vehicle number. To address the issue of data security and alteration of sensitive data blockchain is used so that any alteration can be monitored. As the complete information process is dependent on how correctly the vehicle number is identified, so the number plate recognition module is tested thoroughly under various conditions. Finally user feedback is taken and analyzed to evaluate the feasibility and usability of the proposed application.

[...] Read more.
Cuisine Detection Using the Convolutional Neural Network

By Dipti Pawade Ashwini Dalvi Irfan Siddavatam Myron Carvalho Prajwal Kotian Hima George

DOI: https://doi.org/10.5815/ijeme.2020.03.01, Pub. Date: 8 Jun. 2020

In today’s fast world, everyone wants the information in one click. The same rule applies when you have some food items in front of you. In social events, few cuisines are known to us while some are not. Also, in a few cases, we know the cuisine name, but we are not aware of its nutritional value. This motivated us to develop a system that can identify the cuisine name from the image and gives the nutrient value for the same. Here Convolutional Neural Network (CNN) is used to predict the cuisine name present in an image and then further its nutritional value is calculated based on the information present in a database. User needs to click the image of the cuisine; the application will identify the cuisine name and its nutrition value for standard serving amount considering the cuisine is prepared using the standard recipe.

[...] Read more.
Story Scrambler - Automatic Text Generation Using Word Level RNN-LSTM

By Dipti Pawade Avani Sakhapara Mansi Jain Neha Jain Krushi Gada

DOI: https://doi.org/10.5815/ijitcs.2018.06.05, Pub. Date: 8 Jun. 2018

With the advent of artificial intelligence, the way technology can assist humans is completely revived. Ranging from finance and medicine to music, gaming, and various other domains, it has slowly become an intricate part of our lives. A neural network, a computer system modeled on the human brain, is one of the methods of implementing artificial intelligence. In this paper, we have implemented a recurrent neural network methodology based text generation system called Story Scrambler. Our system aims to generate a new story based on a series of inputted stories.  For new story generation, we have considered two possibilities with respect to nature of inputted stories. Firstly, we have considered the stories with different storyline and characters. Secondly, we have worked with different volumes of the same stories where the storyline is in context with each other and characters are also similar. Results generated by the system are analyzed based on parameters like grammar correctness, linkage of events, interest level and uniqueness.

[...] Read more.
Distributed Ledger Management for an Organization using Blockchains

By Dipti Pawade Sagar Jape Rahul Balasubramanian Mihir Kulkarni Avani Sakhapara

DOI: https://doi.org/10.5815/ijeme.2018.03.01, Pub. Date: 8 May 2018

In the financial systems of the modern era, trust has always been a missing entity; concentrations of power and trust have given birth to numerous breakdowns. To resolve the problems at this end, the paper attempts for a solution using Blockchains, a data structure that allows for the creation of cryptically secured distributed tamperproof ledgers. The paper discusses the data structure and further disserts on the case study in consideration.

[...] Read more.
Augmented Reality Based Campus Guide Application Using Feature Points Object Detection

By Dipti Pawade Avani Sakhapara Maheshwar Mundhe Aniruddha Kamath Devansh Dave

DOI: https://doi.org/10.5815/ijitcs.2018.05.08, Pub. Date: 8 May 2018

These days though GPS is a very common navigation system, still using GPS for everybody is not possible due to the lack of technological awareness. So people follow the traditional method of asking location to the local people around. After taking external guidance, even if one reaches the destination place, one is not able to understand the significance of that place. In this paper, we have discussed an implementation of a Mobile Augmented Reality based application called “ARCampusGo". With this application, one has to just scan the structure/monument to view the details about it. Along with it, this application also provides the names of nearby structure/monuments. One can select any of them and then the route to the selected monument/structure from the current location is rendered. This application renders rich, easy and interactive visual experience to the user. The performance and usability of “ARCampusGo" are evaluated during different daytime and nighttime with the different number of users. The user experience and feedback are considered for performance measurement and enhancement.

[...] Read more.
Other Articles