ISSN: 2074-9074 (Print)
ISSN: 2074-9082 (Online)
DOI: https://doi.org/10.5815/ijigsp
Website: https://www.mecs-press.org/ijigsp
Published By: MECS Press
Frequency: 6 issues per year
Number(s) Available: 145
IJIGSP is committed to bridge the theory and practice of images, graphics, and signal processing. From innovative ideas to specific algorithms and full system implementations, IJIGSP publishes original, peer-reviewed, and high quality articles in the areas of images, graphics, and signal processing. IJIGSP is a well-indexed scholarly journal and is indispensable reading and references for people working at the cutting edge of images, graphics, and signal processing applications.
IJIGSP has been abstracted or indexed by several world class databases: Scopus, SCImago, Google Scholar, CrossRef, Baidu Wenku, IndexCopernicus, IET Inspec, EBSCO, JournalSeek, ULRICH's Periodicals Directory, WorldCat, Academic Journals Database, Stanford University Libraries, Cornell University Library, UniSA Library, CNKI Scholar, ProQuest, J-Gate, ZDB, BASE, OhioLINK, iThenticate, Open Access Articles, Open Science Directory, National Science Library of Chinese Academy of Sciences, The HKU Scholars Hub, etc..
IJIGSP Vol. 18, No. 5, Oct. 2026
REGULAR PAPERS
This study proposes a deep learning-based framework for the automated detection of asthma and chronic obstructive pulmonary disease (COPD) using respiratory sound analysis. Breath sound recordings are preprocessed through resampling, silence trimming, and Wiener filtering to enhance signal quality. Short-Time Fourier Transform (STFT) is employed to convert audio signals into spectrogram representations, which are further augmented using time masking, frequency masking, and time warping to improve model generalization. The proposed model utilizes EfficientNet with multi-scale feature fusion to capture both local and global patterns in respiratory sounds. Experimental results demonstrate that the proposed approach achieves superior performance, with an accuracy of 99.21%, sensitivity of 99.56%, and specificity of 100%, outperforming existing CNN, ResNet, and SVM-based methods. The findings indicate that the proposed method is a reliable and efficient tool for non-invasive respiratory disease diagnosis.
[...] Read more.Tuberculosis (TB) and lung cancer remain leading causes of mortality worldwide, emphasizing the need for reliable automated diagnostic systems. Existing deep learning approaches typically address segmentation and classification as independent tasks or rely on loosely coupled hybrid architectures, limiting joint optimization and interpretability. To address these limitations, this work proposes a Dual-Path ConvNet–Vision Transformer (ViT) hybrid framework for simultaneous pulmonary disease classification and lesion segmentation. Unlike fully shared multi-task models, the proposed design integrates convolutional feature extraction for classification with transformer-based global context modeling for segmentation, followed by feature-level fusion to enhance diagnostic consistency. The framework is evaluated on the IQ-OTH/NCCD dataset, achieving an accuracy of 0.950, recall of 0.951, precision of 0.947, specificity of 0.975, F1-score of 0.949, and AUC of 0.963. Results demonstrate that the proposed hybrid approach provides robust and interpretable performance for pulmonary disease co-analysis while maintaining architectural flexibility.
[...] Read more.Federated learning enables industrial sites to train shared condition-monitoring models without exchanging raw vibration, acoustic, or sensor data, but its openness admits Byzantine clients that submit corrupted updates. Existing robust aggregation rules—Krum, trimmed mean, median, and their variants—operate in Euclidean parameter space and implicitly assume that a malicious update must be far from honest updates to do harm. We show that this assumption fails for models with Kolmogorov–Arnold (KAN) spline heads: an adversary can hold the spline coefficients close to the honest consensus in the L2 sense while distorting the functional shape the spline encodes, an attack we term functional spline-poisoning. We propose ByzaSpline-Fed, a Byzantine-robust aggregation rule that scores clients by Gram-weighted functional B-spline distance on a consensus knot grid and rejects functional outliers. The rule reduces exactly to spline-aware federated averaging under honest participation, admits a breakdown bound of f < K/2 under the standard honest-majority assumption, and supports optional differential privacy. We evaluate across four industrial benchmarks spanning classification, unsupervised acoustic anomaly detection, and remaining-useful-life regression—CWRU, Paderborn, MIMII, and C-MAPSS—reporting task-correct metrics under clean and adversarial conditions. Under thirty percent functional poisoning, Euclidean defenses lose up to thirty-four accuracy points while ByzaSpline-Fed remains within roughly three points of its clean performance. All reported results are averaged over eight independent random seeds, and the improvements over the strongest competing defense are statistically significant under the Wilcoxon signed-rank test (p < 0.05).
[...] Read more.Electrocardiogram (ECG)-based gender identification, which utilizes the electrical activity of the heart, has emerged as a promising approach in biometric and healthcare applications. This study introduces DeepFusion-CNN, a context-aware fusion framework that integrates VGG-19, DenseNet-121, and ResNet-152 using a validation-driven adaptive weighting strategy to improve gender classification performance. Unlike conventional ensemble approaches that use a static averaging strategy, the proposed approach adaptively adjusts each sub-model's contribution based on its validation performance, enabling improved feature representation and classification robustness. This adaptive fusion mechanism allows better-performing models to contribute more significantly, leading to improved overall prediction accuracy compared to individual models and static fusion strategies. The ECG signals are preprocessed using band-pass filtering, followed by R-peak identification using the Pan–Tompkins algorithm. The processed signals are then segmented and converted into 225×225×3 two-dimensional images, making them suitable for transfer learning with pre-trained convolutional models. To maintain a fair evaluation, the data is partitioned on a subject basis before any augmentation, and augmentation is restricted to the training portion only. The framework is evaluated on the PTB and CYBHi datasets, achieving accuracies of 99.08% and 99.13%, respectively. Ablation test results indicate that the feature quality and classification performance are improved after preprocessing and the context-aware fusion strategy. The proposed framework shows strong potential for ECG-based gender classification and could serve as a useful foundation for future advancements in biometric systems and healthcare applications.
[...] Read more.Accurate IoT device classification is essential for secure and efficient network management in heterogeneous environments. However, existing approaches struggle with overlapping traffic patterns and limited generalization across dynamic device behaviours. This paper proposes a Metadata Fusion with Attention-based BiLSTM (MFA-BiLSTM) framework that integrates time-domain and statistical features using an adaptive attention mechanism. The model captures both sequential dependencies and distributional characteristics of network traffic, enhancing feature representation and classification robustness. Experiments conducted on the CIC-IoT-Dataset2022 demonstrate improved performance with an accuracy of 99.40%, precision of 99.39%, recall of 99.36%, and F1-score of 99.38%. The results indicate that the proposed approach achieves reliable and scalable IoT device classification while maintaining interpretability.
[...] Read more.The need for flexible software receivers for research and experimentation has naturally grown due to the rapid development of Global/Regional Navigation Satellite System (GNSS/RNSS) constellations, created to cater to the navigation requirements of various countries. The two primary functions of all such receivers are signal acquisition and signal tracking. One must have a firm understanding of these functions to design or improve an existing navigation receiver. Signal acquisition is the first operation at the receiver, after filtering the received signal to remove noise. It provides the preliminary estimates of the carrier frequency and code phase and assists in identifying the satellites that are in view of the receiver. The accuracy of these estimates is then progressively improved by tracking. To accomplish this refinement using a correlation operation to compare the received signal and locally generated code and carrier replicas, the tracking process uses a tracking loop whose main part is a Phase-Locked Loop (PLL). This step is especially crucial because it eliminates the modulation that was added at the transmitter, enabling the receiver to retrieve the navigational information required for position calculation. The PLL continuously modifies the locally generated code and carrier replicas based on the discriminator output to maintain alignment with the incoming signal. The correlation peaks, when the local replicas and the received signal are well synchronized, showing the effective cancellation of the code and carrier modulation. This makes it possible to reliably extract the navigation message. Using GPS as the reference model, we present in this work a thorough explanation of the tracking mechanism, the structure of the PLL, and the mathematical principles governing its operation. The Indian Regional Navigation Satellite System (IRNSS) signal structure is then subjected to the same methodology. To validate the method we have used the actual intermediate-frequency data from an IRNSS-User Receiver (IRNSS-UR) installed by ISRO at the IRNSS lab of Jain University, Bengaluru. The obtained values of frequency and phase jitter to measure the performance of the tracking loop show that the tracking loop is stable.
[...] Read more.To accurately segment brain tumors and grade gliomas using multi-modal MRI data, MRI-Glioma Net was built as a new 3D Res-UNet framework. The model uses residual learning, multi-scale feature extraction, and attention-enhanced fusion to delineate three heterogeneous tumor subregions: whole tumor (WT), tumor core (TC), and enhancing tumor (ET). It leverages T1, T2, FLAIR, and T1ce sequences. It uses fused latent features to incorporate a specialized classification head for glioma grading (low-grade vs. high-grade), which improves discriminative capability. A glioma dataset with extensive cross-validation was used for evaluation across multiple institutions. MRI-Glioma Net outperformed baseline models such as 3D UNet and Res UNet, achieving Dice Similarity Coefficients (DSCs) of 0.94 (WT), 0.90 (TC), and 0.88 (ET), respectively. An IoU of 0.86, HD95 of 5.4 mm, and a volumetric similarity of 0.95 were all recorded by the model. In terms of grading, it achieved better results than attention UNet and nn UNet, with 95% accuracy, 94% precision, 93% recall, and 93.5% F1-score. With a 7.2 GB GPU usage, a 2.1s inference time per volume, and only a 0.9% accuracy loss post-quantization, the model's efficiency metrics demonstrate its lightweight deployment potential. Additionally, 95% confidence intervals are computed for key metrics to reflect variability across folds. Statistical significance of improvements over baseline models (3D U-Net, Res-UNet, Attention U-Net, nnU-Net) is evaluated using paired statistical tests (e.g., Wilcoxon signed-rank test), confirming that performance gains are not due to random variation. Efficiency is further validated using inference time per volume, GPU memory consumption, and post-quantization performance, ensuring practical deployment feasibility. The anatomical fidelity is shown by qualitative overlays to be superior, and interpretability is improved by Grad-CAM and error maps. MRI-Glioma Net offers a feasible, effective, and interpretable way to classify gliomas and diagnose tumors in real time in clinical settings. MRI-Glioma Net represents a major step forward in neuro-oncology imaging; its strong performance across segmentation and grading tasks suggests it could be useful for pre-surgical planning, prognosis, and monitoring treatment response.
[...] Read more.Traumatic injuries in emergency response scenarios require rapid and reliable assessment methods capable of supporting medical prioritisation and efficient resource management. This paper presents a multispectral edge deep learning system designed for early injury detection, assistance prioritisation, and financial-resource decision support in emergency response environments. The proposed approach integrates RGB, thermal, near-infrared (NIR), depth, and contextual information within an edge-based analytical framework that combines multispectral data fusion, adaptive attention-based weighting of sensor modalities, convolutional neural networks for feature extraction and lesion segmentation, severity classification, triage index estimation, and resource allocation optimisation. The developed mathematical model describes the interaction among sensory processing, local inference, decision-making, and feedback-based model updating, and system stability is analysed using a Lyapunov-based approach under specified modelling assumptions. The system was evaluated on a simulated dataset containing N = 1000 multispectral observations, including 150 test samples for final performance assessment. The obtained results demonstrated promising performance on simulated data, achieving Accuracy = 0.912 ± 0.015, Recallcrit = 0.941 ± 0.021, Precisioncrit = 0.876 ± 0.024, and F1-score = 0.908 ± 0.018. The average edge inference delay was 425 ms, satisfying the considered real-time processing threshold of 500 ms. Compared with the baseline emergency response scenario without multispectral edge-based decision support, the proposed approach reduced triage decision errors by 18.6% and unproductive resource utilisation costs by 16.4%. The contribution of this research is the formulation of an integrated multispectral edge deep learning framework that combines injury detection, automated triage assessment, resource optimisation, and financial risk considerations within a unified decision-support architecture. The results indicate the potential of the proposed approach to improve emergency response processes. However, further validation on real-world clinical and operational datasets is required to confirm its practical applicability.
[...] Read more.Brain tumor refers to the abnormal enhancement of a collection of tissues or cells within the brain, occurring when brain cells grow improperly. Brain Tumors have currently been recognized as a major cause of mortality worldwide. As the most fatal category of cancer, brain tumors make detection and treatment critical for protecting lives. Single-Photon Emission Computed Tomography (SPECT) provides important imaging insights for detecting and evaluating brain cancer. However, variations in medical scenarios and imaging techniques make it difficult to improve models that are universally applicable across domains. This research suggests a novel framework for improving SPECT-based brain tumor diagnosis by combining Domain Adversarial Neural Networks (DANN) with Model Agnostic Meta Learning (MAML) and anisotropic diffusion filtering. The anisotropic diffusion filter reduces noise while maintaining tumor relevant structures, thereby preserving critical spatial properties for analysis. DANN supports domain adaptation by learning domain-invariant representations, allowing the model to perform consistently across different datasets. MAML enhances the model’s capability to adapt to new, unfamiliar domains with insufficient data, hence boosting its therapeutic value. Experimental results of the proposed DANN-MAML model with anisotropic diffusion filtering over diverse SPECT datasets outperformed existing strategies regarding both accuracy and flexibility, making it a promising tool for brain tumor identification.
[...] Read more.This paper presents Block-Based Compressive Sensing (BCS) of images using both traditional and learning-based approaches. An Image Block Mapping Lemma, formulated as a modified version of the Johnson-Lindenstrauss Lemma for image data, is proposed in this work, which justifies distance preservation between image blocks during mapping from the original domain to the compressed domain in block-based compressive sensing. The paper provides both theoretical proof and experimental validation of the proposed lemma. To quantitatively analyse distance preservation between image blocks in the original and compressed domains during block-based compressive sensing using traditional approaches, a new metric termed Distance Ratio (DR) is introduced. The difficulty of obtaining accurate real-time reconstruction using conventional block-based compressive sensing methods has encouraged the transition toward adaptive deep learning approaches. For learning-based sensing and reconstruction, the existing AutoBCS framework is enhanced by introducing an Initial Reconstruction Refinement Network (IRR-Net) between the initial and final reconstruction stages. Using the proposed model, improved reconstruction quality is achieved, with average PSNR gains of 0.69-1.62 dB and SSIM improvements of 0.01-0.02 for sampling rates of 0.30, 0.25, 0.10, and 0.04 across multiple benchmark datasets compared with the baseline architecture. The proposed model introduces an additional 0.05 million parameters and increases the computational cost by 3.64 GFLOPs due to the residual refinement module, while reducing the reconstruction time compared with the baseline AutoBCS framework. Experimental results demonstrate that the proposed model exhibits improved preservation of complex structures, edges, and fine details, particularly for images containing rich textures, dense structures, and high-frequency content.
[...] Read more.Adsorption of dyes present in textile industry waste water is an important process for environmental remediation and pollution control. Dyes are complicated organic compounds used in textile dyeing processes. The presence of dyes in wastewater is of major environmental concern because of their recalcitrance and potential toxicity. Adsorption technology is a promising solution for the removal of dyes from textile wastewater. It uses adsorbent materials to capture and immobilize dye molecules from aqueous solution. Machine learning techniques have been used to reduce the analytical efforts for prediction of concentration of dye solution in parts per million using a UV spectrophotometer.
In the research work, a prediction technique using Machine Learning and image processing is used to identify the color of images of samples and corresponding adsorption of the dye. The suggested model has achieved correct dye concentration estimates using retrieved image attributes with 93% accuracy with experimental data.
This paper introduces BlinkFusion, a real-time, interpretable, and detector-independent pipeline for analyzing blink events. The system finds eye ROIs using YOLOv5 as the main detector and a Haar cascade as a backup that has been calibrated. It then stabilizes the detections using lightweight tracking. A small landmark regressor inside each ROI gives six points to calculate the Eye Aspect Ratio (EAR), which keeps the geometric meaning. An uncertainty-weighted smoother combines pose, landmark, and detector confidence. An online hysteresis state machine with minimum-duration and derivative gates makes blink onsets and offsets. Platt-calibrated and fused detector confidences make strong arbitration possible at a reasonable cost. Performance is assessed on EyeBlinkDB (RGB ≥25 fps, ~720p), using subject-independent 10-fold splits and temporal-IoU matching. The model achieves Precision 0.928 ± 0.008, Recall 0.945 ± 0.008, F1 0.936 ± 0.008, AP@tIoU=0.3 = 0.962 ± 0.006, with timing precision of 28.1 ± 1.7 ms onset MAE and 37.1 ± 2.3 ms offset MAE. Condition stratification validates robustness: The values for Bright/Dim F1 are 0.947/0.922, for Frontal/±15°/≥30° yaw F1 they are 0.952/0.936/0.903, for Glasses (No/Yes) they are 0.946/0.925, and for Occlusion (No/Yes) they are 0.948/0.906. With adaptive scheduling (YOLO invocation rate ρ≈0.045), pipeline runs at about 32 FPS on the CPU (about 90 FPS on the observed frame loop) with a fixed 3-frame delay, which is fast enough for real-time use on cheap hardware. Ablations demonstrate progressive improvements resulting from calibrated fusion, tracking, quality-weighted EMA, and hysteresis (F1: 0.903 → 0.952). The approach is modular (you may switch the detector), doesn't need bounding boxes, and makes judgments that can be explained using EAR traces and thresholds. So, BlinkFusion is a useful, ready-to-use solution for HCI and clinical contexts that need clear, accurate, and quick blink analytics.
[...] Read more.Mushrooms are the most familiar delicious food which is cholesterol free as well as rich in vitamins and minerals. Though nearly 45,000 species of mushrooms have been known throughout the world, most of them are poisonous and few are lethally poisonous. Identifying edible or poisonous mushroom through the naked eye is quite difficult. Even there is no easy rule for edibility identification using machine learning methods that work for all types of data. Our aim is to find a robust method for identifying mushrooms edibility with better performance than existing works. In this paper, three ensemble methods are used to detect the edibility of mushrooms: Bagging, Boosting, and random forest. By using the most significant features, five feature sets are made for making five base models of each ensemble method. The accuracy is measured for ensemble methods using five both fixed feature set-based models and randomly selected feature set based models, for two types of test sets. The result shows that better performance is obtained for methods made of fixed feature sets-based models than randomly selected feature set-based models. The highest accuracy is obtained for the proposed model-based random forest for both test sets.
[...] Read more.This paper presents a design and development of an Artificial Intelligence (AI) based mobile application to detect the type of skin disease. Skin diseases are a serious hazard to everyone throughout the world. However, it is difficult to make accurate skin diseases diagnosis. In this work, Deep learning algorithms Convolution Neural Networks (CNN) is proposed to classify skin diseases on the HAM10000 dataset. An extensive review of research articles on object identification methods and a comparison of their relative qualities were given to find a method that would work well for detecting skin diseases. The CNN-based technique was recognized as the best method for identifying skin diseases. A mobile application, on the other hand, is built for quick and accurate action. By looking at an image of the afflicted area at the beginning of a skin illness, it assists patients and dermatologists in determining the kind of disease present. Its resilience in detecting the impacted region considerably faster with nearly 2x fewer computations than the standard MobileNet model results in low computing efforts. This study revealed that MobileNet with transfer learning yielding an accuracy of about 85% is the most suitable model for automatic skin disease identification. According to these findings, the suggested approach can assist general practitioners in quickly and accurately diagnosing skin diseases using the smart phone.
[...] Read more.Classifying and predicting banana shelf life is vital for optimizing storage and distribution in agriculture. Traditional methods, relying on subjective visual inspection, are inconsistent and time-intensive. This study presents a new, non-destructive approach combining thermal imaging, and machine learning to classify naturally ripened and artificially ripened bananas and forecast their shelf life. Preprocessed thermal images are flattened, segmented into fixed-size patches, and then linearly projected into feature tokens. Position embeddings are incorporated to retain spatial information, and the sequence is processed by a Vision Transformer (ViT) encoder, which leverages self-attention mechanisms to model relationships between patches. The [CLS] token output is subsequently processed through fully connected layers for final classification, achieving 97.59% accuracy. Validation using t-SNE visualization demonstrated clear class separability, and receiver operating characteristic (ROC) curves confirmed robust performance. With an MSE of 0.10, MAE of 0.18, and R2 score of 0.85, the random forest algorithm performed exceptionally well at predicting the shelf life of artificially ripened bananas. This approach offers significant advantages, including improved accuracy, reduced subjectivity, and efficiency in data processing. By integrating thermal imaging with advanced models, the proposed method enhances agricultural supply chain management and promotes precision in ripening classification and shelf life prediction.
[...] Read more.In the field of medical image analysis, supervised deep learning strategies have achieved significant development, while these methods rely on large labeled datasets. Self-Supervised learning (SSL) provides a new strategy to pre-train a neural network with unlabeled data. This is a new unsupervised learning paradigm that has achieved significant breakthroughs in recent years. So, more and more researchers are trying to utilize SSL methods for medical image analysis, to meet the challenge of assembling large medical datasets. To our knowledge, so far there still a shortage of reviews of self-supervised learning methods in the field of medical image analysis, our work of this article aims to fill this gap and comprehensively review the application of self-supervised learning in the medical field. This article provides the latest and most detailed overview of self-supervised learning in the medical field and promotes the development of unsupervised learning in the field of medical imaging. These methods are divided into three categories: context-based, generation-based, and contrast-based, and then show the pros and cons of each category and evaluates their performance in downstream tasks. Finally, we conclude with the limitations of the current methods and discussed the future direction.
[...] Read more.Image Processing is the art of examining, identifying and judging the significances of the Images. Image enhancement refers to attenuation, or sharpening, of image features such as edgels, boundaries, or contrast to make the processed image more useful for analysis. Image enhancement procedures utilize the computers to provide good and improved images for study by the human interpreters. In this paper we proposed a novel method that uses the Genetic Algorithm with Multi-objective criteria to find more enhance version of images. The proposed method has been verified with benchmark images in Image Enhancement. The simple Genetic Algorithm may not explore much enough to find out more enhanced image. In the proposed method three objectives are taken in to consideration. They are intensity, entropy and number of edgels. Proposed algorithm achieved automatic image enhancement criteria by incorporating the objectives (intensity, entropy, edges). We review some of the existing Image Enhancement technique. We also compared the results of our algorithms with another Genetic Algorithm based techniques. We expect that further improvements can be achieved by incorporating linear relationship between some other techniques.
[...] Read more.Image analysis belongs to the area of computer vision and pattern recognition. These areas are also a part of digital image processing, where researchers have a great attention in the area of content retrieval information from various types of images having complex background, low contrast background or multi-spectral background etc. These contents may be found in any form like texture data, shape, and objects. Text Region Extraction as a content from an mage is a class of problems in Digital Image Processing Applications that aims to provides necessary information which are widely used in many fields medical imaging, pattern recognition, Robotics, Artificial intelligent Transport systems etc. To extract the text data information has becomes a challenging task. Since, Text extraction are very useful for identifying and analysis the whole information about image, Therefore, In this paper, we propose a unified framework by combining morphological operations and Genetic Algorithms for extracting and analyzing the text data region which may be embedded in an image by means of variety of texts: font, size, skew angle, distortion by slant and tilt, shape of the object which texts are on, etc. We have established our proposed methods on gray level image sets and make qualitative and quantitative comparisons with other existing methods and concluded that proposed method is better than others.
[...] Read more.This paper performs three different contrast testing methods, namely contrast stretching, histogram equalization, and CLAHE using a median filter. Poor quality images will be corrected and performed with a median filter removal filter. STARE dataset images that use images with different contrast values for each image. For this reason, evaluating the results of the three parameters tested are; MSE, PSNR, and SSIM. With the gray level scale image and contrast stretching which stretches the pixel value by stretching the stretchlim technique with the MSE result are 9.15, PSNR is 42.14 dB, and SSIM is 0.88. And the HE method and median filter with the results of the average value of MSE is 18.67, PSNR is 41.33 dB, and SSIM is 0.77. Whereas for CLAHE and median filters the average yield of MSE is 28.42, PSNR is 35.30 dB, and SSIM is 0.86. From the test results, it can be seen that the proposed method has MSE and PSNR values as well as SSIM values.
[...] Read more.Denoising is a vital aspect of image preprocessing, often explored to eliminate noise in an image to restore its proper characteristic formation and clarity. Unfortunately, noise often degrades the quality of valuable images, making them meaningless for practical applications. Several methods have been deployed to address this problem, but the quality of the recovered images still requires enhancement for efficient applications in practice. In this paper, a wavelet-based universal thresholding technique that possesses the capacity to optimally denoise highly degraded noisy images with both uniform and non-uniform variations in illumination and contrast is proposed. The proposed method, herein referred to as the modified wavelet-based universal thresholding (MWUT), compared to three state-of-the-art denoising techniques, was employed to denoise five noisy images. In order to appraise the qualities of the images obtained, seven performance indicators comprising the Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Structural Content (SC), Peak Signal to Noise Ratio (PSNR), Structural Similarity Index Method (SSIM), Signal-to-Reconstruction-Error Ratio (SRER), Blind Spatial Quality Evaluator (NIQE), and Blind/Referenceless Image Spatial Quality Evaluator (BRISQUE) were employed. The first five indicators – RMSE, MAE, SC, PSNR, SSIM, and SRER- are reference indicators, while the remaining two – NIQE and BRISQUE- are referenceless. For the superior performance of the proposed wavelet threshold algorithm, the SC, PSNR, SSIM, and SRER must be higher, while lower values of NIQE, BRISQUE, RMSE, and MAE are preferred. A higher and better value of PSNR, SSIM, and SRER in the final results shows the superior performance of our proposed MWUT denoising technique over the preliminaries. Lower NIQE, BRISQUE, RMSE, and MAE values also indicate higher and better image quality results using the proposed modified wavelet-based universal thresholding technique over the existing schemes. The modified wavelet-based universal thresholding technique would find practical applications in digital image processing and enhancement.
[...] Read more.Image segmentation plays a crucial role in effective understanding of digital images. Past few decades saw hundreds of research contributions in this field. However, the research on the existence of general purpose segmentation algorithm that suits for variety of applications is still very much active. Among the many approaches in performing image segmentation, graph based approach is gaining popularity primarily due to its ability in reflecting global image properties. This paper critically reviews existing important graph based segmentation methods. The review is done based on the classification of various segmentation algorithms within the framework of graph based approaches. The major four categorizations we have employed for the purpose of review are: graph cut based methods, interactive methods, minimum spanning tree based methods and pyramid based methods. This review not only reveals the pros in each method and category but also explores its limitations. In addition, the review highlights the need for creating a database for benchmarking intensity based algorithms, and the need for further research in graph based segmentation for automated real time applications.
[...] Read more.During past few years, brain tumor segmentation in magnetic resonance imaging (MRI) has become an emergent research area in the ?eld of medical imaging system. Brain tumor detection helps in finding the exact size and location of tumor. An efficient algorithm is proposed in this paper for tumor detection based on segmentation and morphological operators. Firstly quality of scanned image is enhanced and then morphological operators are applied to detect the tumor in the scanned image.
[...] Read more.Mushrooms are the most familiar delicious food which is cholesterol free as well as rich in vitamins and minerals. Though nearly 45,000 species of mushrooms have been known throughout the world, most of them are poisonous and few are lethally poisonous. Identifying edible or poisonous mushroom through the naked eye is quite difficult. Even there is no easy rule for edibility identification using machine learning methods that work for all types of data. Our aim is to find a robust method for identifying mushrooms edibility with better performance than existing works. In this paper, three ensemble methods are used to detect the edibility of mushrooms: Bagging, Boosting, and random forest. By using the most significant features, five feature sets are made for making five base models of each ensemble method. The accuracy is measured for ensemble methods using five both fixed feature set-based models and randomly selected feature set based models, for two types of test sets. The result shows that better performance is obtained for methods made of fixed feature sets-based models than randomly selected feature set-based models. The highest accuracy is obtained for the proposed model-based random forest for both test sets.
[...] Read more.Image Processing is the art of examining, identifying and judging the significances of the Images. Image enhancement refers to attenuation, or sharpening, of image features such as edgels, boundaries, or contrast to make the processed image more useful for analysis. Image enhancement procedures utilize the computers to provide good and improved images for study by the human interpreters. In this paper we proposed a novel method that uses the Genetic Algorithm with Multi-objective criteria to find more enhance version of images. The proposed method has been verified with benchmark images in Image Enhancement. The simple Genetic Algorithm may not explore much enough to find out more enhanced image. In the proposed method three objectives are taken in to consideration. They are intensity, entropy and number of edgels. Proposed algorithm achieved automatic image enhancement criteria by incorporating the objectives (intensity, entropy, edges). We review some of the existing Image Enhancement technique. We also compared the results of our algorithms with another Genetic Algorithm based techniques. We expect that further improvements can be achieved by incorporating linear relationship between some other techniques.
[...] Read more.Denoising is a vital aspect of image preprocessing, often explored to eliminate noise in an image to restore its proper characteristic formation and clarity. Unfortunately, noise often degrades the quality of valuable images, making them meaningless for practical applications. Several methods have been deployed to address this problem, but the quality of the recovered images still requires enhancement for efficient applications in practice. In this paper, a wavelet-based universal thresholding technique that possesses the capacity to optimally denoise highly degraded noisy images with both uniform and non-uniform variations in illumination and contrast is proposed. The proposed method, herein referred to as the modified wavelet-based universal thresholding (MWUT), compared to three state-of-the-art denoising techniques, was employed to denoise five noisy images. In order to appraise the qualities of the images obtained, seven performance indicators comprising the Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Structural Content (SC), Peak Signal to Noise Ratio (PSNR), Structural Similarity Index Method (SSIM), Signal-to-Reconstruction-Error Ratio (SRER), Blind Spatial Quality Evaluator (NIQE), and Blind/Referenceless Image Spatial Quality Evaluator (BRISQUE) were employed. The first five indicators – RMSE, MAE, SC, PSNR, SSIM, and SRER- are reference indicators, while the remaining two – NIQE and BRISQUE- are referenceless. For the superior performance of the proposed wavelet threshold algorithm, the SC, PSNR, SSIM, and SRER must be higher, while lower values of NIQE, BRISQUE, RMSE, and MAE are preferred. A higher and better value of PSNR, SSIM, and SRER in the final results shows the superior performance of our proposed MWUT denoising technique over the preliminaries. Lower NIQE, BRISQUE, RMSE, and MAE values also indicate higher and better image quality results using the proposed modified wavelet-based universal thresholding technique over the existing schemes. The modified wavelet-based universal thresholding technique would find practical applications in digital image processing and enhancement.
[...] Read more.Ultrasound based breast screening is gaining attention recently especially for dense breast. The technological advancement, cancer awareness, and cost-safety-availability benefits lead rapid rise of breast ultrasound market. The irregular shape, intensity variation, and additional blood vessels of malignant cancer are distinguishable in ultrasound images from the benign phase. However, classification of breast cancer using ultrasound images is a difficult process owing to speckle noise and complex textures of breast. In this paper, a breast cancer classification method is presented using VGG16 model based transfer learning approach. We have used median filter to despeckle the images. The layers for convolution process of the pretrained VGG16 model along with the maxpooling layers have been used as feature extractor and a proposed fully connected two layers deep neural network has been designed as classifier. Adam optimizer is used with learning rate of 0.001 and binary cross-entropy is chosen as the loss function for model optimization. Dropout of hidden layers is used to avoid overfitting. Breast Ultrasound images from two databases (total 897 images) have been combined to train, validate and test the performance and generalization strength of the classifier. Experimental results showed the training accuracy as 98.2% and testing accuracy as 91% for blind testing data with a reduced of computational complexity. Gradient class activation mapping (Grad-CAM) technique has been used to visualize and check the targeted regions localization effort at the final convolutional layer and found as noteworthy. The outcomes of this work might be useful for the clinical applications of breast cancer diagnosis.
[...] Read more.This paper presents a design and development of an Artificial Intelligence (AI) based mobile application to detect the type of skin disease. Skin diseases are a serious hazard to everyone throughout the world. However, it is difficult to make accurate skin diseases diagnosis. In this work, Deep learning algorithms Convolution Neural Networks (CNN) is proposed to classify skin diseases on the HAM10000 dataset. An extensive review of research articles on object identification methods and a comparison of their relative qualities were given to find a method that would work well for detecting skin diseases. The CNN-based technique was recognized as the best method for identifying skin diseases. A mobile application, on the other hand, is built for quick and accurate action. By looking at an image of the afflicted area at the beginning of a skin illness, it assists patients and dermatologists in determining the kind of disease present. Its resilience in detecting the impacted region considerably faster with nearly 2x fewer computations than the standard MobileNet model results in low computing efforts. This study revealed that MobileNet with transfer learning yielding an accuracy of about 85% is the most suitable model for automatic skin disease identification. According to these findings, the suggested approach can assist general practitioners in quickly and accurately diagnosing skin diseases using the smart phone.
[...] Read more.In the field of medical image analysis, supervised deep learning strategies have achieved significant development, while these methods rely on large labeled datasets. Self-Supervised learning (SSL) provides a new strategy to pre-train a neural network with unlabeled data. This is a new unsupervised learning paradigm that has achieved significant breakthroughs in recent years. So, more and more researchers are trying to utilize SSL methods for medical image analysis, to meet the challenge of assembling large medical datasets. To our knowledge, so far there still a shortage of reviews of self-supervised learning methods in the field of medical image analysis, our work of this article aims to fill this gap and comprehensively review the application of self-supervised learning in the medical field. This article provides the latest and most detailed overview of self-supervised learning in the medical field and promotes the development of unsupervised learning in the field of medical imaging. These methods are divided into three categories: context-based, generation-based, and contrast-based, and then show the pros and cons of each category and evaluates their performance in downstream tasks. Finally, we conclude with the limitations of the current methods and discussed the future direction.
[...] Read more.Image analysis belongs to the area of computer vision and pattern recognition. These areas are also a part of digital image processing, where researchers have a great attention in the area of content retrieval information from various types of images having complex background, low contrast background or multi-spectral background etc. These contents may be found in any form like texture data, shape, and objects. Text Region Extraction as a content from an mage is a class of problems in Digital Image Processing Applications that aims to provides necessary information which are widely used in many fields medical imaging, pattern recognition, Robotics, Artificial intelligent Transport systems etc. To extract the text data information has becomes a challenging task. Since, Text extraction are very useful for identifying and analysis the whole information about image, Therefore, In this paper, we propose a unified framework by combining morphological operations and Genetic Algorithms for extracting and analyzing the text data region which may be embedded in an image by means of variety of texts: font, size, skew angle, distortion by slant and tilt, shape of the object which texts are on, etc. We have established our proposed methods on gray level image sets and make qualitative and quantitative comparisons with other existing methods and concluded that proposed method is better than others.
[...] Read more.Diabetic retinopathy is one of the most serious eye diseases and can lead to permanent blindness if not diagnosed early. The main cause of this is diabetes. Not every diabetic will develop diabetic retinopathy, but the risk of developing diabetes is undeniable. This requires the early diagnosis of Diabetic retinopathy. Segmentation is one of the approaches which is useful for detecting the blood vessels in the retinal image. This paper proposed the three models based on a deep learning approach for recognizing blood vessels from retinal images using region-based segmentation techniques. The proposed model consists of four steps preprocessing, Augmentation, Model training, and Performance measure. The augmented retinal images are fed to the three models for training and finally, get the segmented image. The proposed three models are applied on publically available data set of DRIVE, STARE, and HRF. It is observed that more thin blood vessels are segmented on the retinal image in the HRF dataset using model-3. The performance of proposed three models is compare with other state-of-art-methods of blood vessels segmentation of DRIVE, STARE, and HRF datasets.
[...] Read more.Image reconstruction is the process of generating an image of an object from the signals captured by the scanning machine. Medical imaging is an interdisciplinary field combining physics, biology, mathematics and computational sciences. This paper provides a complete overview of image reconstruction process in MRI (Magnetic Resonance Imaging). It reviews the computational aspect of medical image reconstruction. MRI is one of the commonly used medical imaging techniques. The data collected by MRI scanner for image reconstruction is called the k-space data. For reconstructing an image from k-space data, there are various algorithms such as Homodyne algorithm, Zero Filling method, Dictionary Learning, and Projections onto Convex Set method. All the characteristics of k-space data and MRI data collection technique are reviewed in detail. The algorithms used for image reconstruction discussed in detail along with their pros and cons. Various modern magnetic resonance imaging techniques like functional MRI, diffusion MRI have also been introduced. The concepts of classical techniques like Expectation Maximization, Sensitive Encoding, Level Set Method, and the recent techniques such as Alternating Minimization, Signal Modeling, and Sphere Shaped Support Vector Machine are also reviewed. It is observed that most of these techniques enhance the gradient encoding and reduce the scanning time. Classical algorithms provide undesirable blurring effect when the degree of phase variation is high in partial k-space. Modern reconstructions algorithms such as Dictionary learning works well even with high phase variation as these are iterative procedures.
[...] Read more.Nowadays, the primary concern of any society is providing safety to an individual. It is very hard to recognize the human behaviour and identify whether it is suspicious or normal. Deep learning approaches paved the way for the development of various machine learning and artificial intelligence. The proposed system detects real-time human activity using a convolutional neural network. The objective of the study is to develop a real-time application for Activity recognition using with and without transfer learning methods. The proposed system considers criminal, suspicious and normal categories of activities. Differentiate suspicious behaviour videos are collected from different peoples(men/women). This proposed system is used to detect suspicious activities of a person. The novel 2D-CNN, pre-trained VGG-16 and ResNet50 is trained on video frames of human activities such as normal and suspicious behaviour. Similarly, the transfer learning in VGG16 and ResNet50 is trained using human suspicious activity datasets. The results show that the novel 2D-CNN, VGG16, and ResNet50 without transfer learning achieve accuracy of 98.96%, 97.84%, and 99.03%, respectively. In Kaggle/real-time video, the proposed system employing 2D-CNN outperforms the pre-trained model VGG16. The trained model is used to classify the activity in the real-time captured video. The performance obtained on ResNet50 with transfer learning accuracy of 99.18% is higher than VGG16 transfer learning accuracy of 98.36%.
[...] Read more.