IJIGSP Vol. 18, No. 5, Oct. 2026
Cover page and Table of Contents: PDF (size: 1048KB)
REGULAR PAPERS
This study proposes a deep learning-based framework for the automated detection of asthma and chronic obstructive pulmonary disease (COPD) using respiratory sound analysis. Breath sound recordings are preprocessed through resampling, silence trimming, and Wiener filtering to enhance signal quality. Short-Time Fourier Transform (STFT) is employed to convert audio signals into spectrogram representations, which are further augmented using time masking, frequency masking, and time warping to improve model generalization. The proposed model utilizes EfficientNet with multi-scale feature fusion to capture both local and global patterns in respiratory sounds. Experimental results demonstrate that the proposed approach achieves superior performance, with an accuracy of 99.21%, sensitivity of 99.56%, and specificity of 100%, outperforming existing CNN, ResNet, and SVM-based methods. The findings indicate that the proposed method is a reliable and efficient tool for non-invasive respiratory disease diagnosis.
[...] Read more.Tuberculosis (TB) and lung cancer remain leading causes of mortality worldwide, emphasizing the need for reliable automated diagnostic systems. Existing deep learning approaches typically address segmentation and classification as independent tasks or rely on loosely coupled hybrid architectures, limiting joint optimization and interpretability. To address these limitations, this work proposes a Dual-Path ConvNet–Vision Transformer (ViT) hybrid framework for simultaneous pulmonary disease classification and lesion segmentation. Unlike fully shared multi-task models, the proposed design integrates convolutional feature extraction for classification with transformer-based global context modeling for segmentation, followed by feature-level fusion to enhance diagnostic consistency. The framework is evaluated on the IQ-OTH/NCCD dataset, achieving an accuracy of 0.950, recall of 0.951, precision of 0.947, specificity of 0.975, F1-score of 0.949, and AUC of 0.963. Results demonstrate that the proposed hybrid approach provides robust and interpretable performance for pulmonary disease co-analysis while maintaining architectural flexibility.
[...] Read more.Federated learning enables industrial sites to train shared condition-monitoring models without exchanging raw vibration, acoustic, or sensor data, but its openness admits Byzantine clients that submit corrupted updates. Existing robust aggregation rules—Krum, trimmed mean, median, and their variants—operate in Euclidean parameter space and implicitly assume that a malicious update must be far from honest updates to do harm. We show that this assumption fails for models with Kolmogorov–Arnold (KAN) spline heads: an adversary can hold the spline coefficients close to the honest consensus in the L2 sense while distorting the functional shape the spline encodes, an attack we term functional spline-poisoning. We propose ByzaSpline-Fed, a Byzantine-robust aggregation rule that scores clients by Gram-weighted functional B-spline distance on a consensus knot grid and rejects functional outliers. The rule reduces exactly to spline-aware federated averaging under honest participation, admits a breakdown bound of f < K/2 under the standard honest-majority assumption, and supports optional differential privacy. We evaluate across four industrial benchmarks spanning classification, unsupervised acoustic anomaly detection, and remaining-useful-life regression—CWRU, Paderborn, MIMII, and C-MAPSS—reporting task-correct metrics under clean and adversarial conditions. Under thirty percent functional poisoning, Euclidean defenses lose up to thirty-four accuracy points while ByzaSpline-Fed remains within roughly three points of its clean performance. All reported results are averaged over eight independent random seeds, and the improvements over the strongest competing defense are statistically significant under the Wilcoxon signed-rank test (p < 0.05).
[...] Read more.Electrocardiogram (ECG)-based gender identification, which utilizes the electrical activity of the heart, has emerged as a promising approach in biometric and healthcare applications. This study introduces DeepFusion-CNN, a context-aware fusion framework that integrates VGG-19, DenseNet-121, and ResNet-152 using a validation-driven adaptive weighting strategy to improve gender classification performance. Unlike conventional ensemble approaches that use a static averaging strategy, the proposed approach adaptively adjusts each sub-model's contribution based on its validation performance, enabling improved feature representation and classification robustness. This adaptive fusion mechanism allows better-performing models to contribute more significantly, leading to improved overall prediction accuracy compared to individual models and static fusion strategies. The ECG signals are preprocessed using band-pass filtering, followed by R-peak identification using the Pan–Tompkins algorithm. The processed signals are then segmented and converted into 225×225×3 two-dimensional images, making them suitable for transfer learning with pre-trained convolutional models. To maintain a fair evaluation, the data is partitioned on a subject basis before any augmentation, and augmentation is restricted to the training portion only. The framework is evaluated on the PTB and CYBHi datasets, achieving accuracies of 99.08% and 99.13%, respectively. Ablation test results indicate that the feature quality and classification performance are improved after preprocessing and the context-aware fusion strategy. The proposed framework shows strong potential for ECG-based gender classification and could serve as a useful foundation for future advancements in biometric systems and healthcare applications.
[...] Read more.Accurate IoT device classification is essential for secure and efficient network management in heterogeneous environments. However, existing approaches struggle with overlapping traffic patterns and limited generalization across dynamic device behaviours. This paper proposes a Metadata Fusion with Attention-based BiLSTM (MFA-BiLSTM) framework that integrates time-domain and statistical features using an adaptive attention mechanism. The model captures both sequential dependencies and distributional characteristics of network traffic, enhancing feature representation and classification robustness. Experiments conducted on the CIC-IoT-Dataset2022 demonstrate improved performance with an accuracy of 99.40%, precision of 99.39%, recall of 99.36%, and F1-score of 99.38%. The results indicate that the proposed approach achieves reliable and scalable IoT device classification while maintaining interpretability.
[...] Read more.The need for flexible software receivers for research and experimentation has naturally grown due to the rapid development of Global/Regional Navigation Satellite System (GNSS/RNSS) constellations, created to cater to the navigation requirements of various countries. The two primary functions of all such receivers are signal acquisition and signal tracking. One must have a firm understanding of these functions to design or improve an existing navigation receiver. Signal acquisition is the first operation at the receiver, after filtering the received signal to remove noise. It provides the preliminary estimates of the carrier frequency and code phase and assists in identifying the satellites that are in view of the receiver. The accuracy of these estimates is then progressively improved by tracking. To accomplish this refinement using a correlation operation to compare the received signal and locally generated code and carrier replicas, the tracking process uses a tracking loop whose main part is a Phase-Locked Loop (PLL). This step is especially crucial because it eliminates the modulation that was added at the transmitter, enabling the receiver to retrieve the navigational information required for position calculation. The PLL continuously modifies the locally generated code and carrier replicas based on the discriminator output to maintain alignment with the incoming signal. The correlation peaks, when the local replicas and the received signal are well synchronized, showing the effective cancellation of the code and carrier modulation. This makes it possible to reliably extract the navigation message. Using GPS as the reference model, we present in this work a thorough explanation of the tracking mechanism, the structure of the PLL, and the mathematical principles governing its operation. The Indian Regional Navigation Satellite System (IRNSS) signal structure is then subjected to the same methodology. To validate the method we have used the actual intermediate-frequency data from an IRNSS-User Receiver (IRNSS-UR) installed by ISRO at the IRNSS lab of Jain University, Bengaluru. The obtained values of frequency and phase jitter to measure the performance of the tracking loop show that the tracking loop is stable.
[...] Read more.To accurately segment brain tumors and grade gliomas using multi-modal MRI data, MRI-Glioma Net was built as a new 3D Res-UNet framework. The model uses residual learning, multi-scale feature extraction, and attention-enhanced fusion to delineate three heterogeneous tumor subregions: whole tumor (WT), tumor core (TC), and enhancing tumor (ET). It leverages T1, T2, FLAIR, and T1ce sequences. It uses fused latent features to incorporate a specialized classification head for glioma grading (low-grade vs. high-grade), which improves discriminative capability. A glioma dataset with extensive cross-validation was used for evaluation across multiple institutions. MRI-Glioma Net outperformed baseline models such as 3D UNet and Res UNet, achieving Dice Similarity Coefficients (DSCs) of 0.94 (WT), 0.90 (TC), and 0.88 (ET), respectively. An IoU of 0.86, HD95 of 5.4 mm, and a volumetric similarity of 0.95 were all recorded by the model. In terms of grading, it achieved better results than attention UNet and nn UNet, with 95% accuracy, 94% precision, 93% recall, and 93.5% F1-score. With a 7.2 GB GPU usage, a 2.1s inference time per volume, and only a 0.9% accuracy loss post-quantization, the model's efficiency metrics demonstrate its lightweight deployment potential. Additionally, 95% confidence intervals are computed for key metrics to reflect variability across folds. Statistical significance of improvements over baseline models (3D U-Net, Res-UNet, Attention U-Net, nnU-Net) is evaluated using paired statistical tests (e.g., Wilcoxon signed-rank test), confirming that performance gains are not due to random variation. Efficiency is further validated using inference time per volume, GPU memory consumption, and post-quantization performance, ensuring practical deployment feasibility. The anatomical fidelity is shown by qualitative overlays to be superior, and interpretability is improved by Grad-CAM and error maps. MRI-Glioma Net offers a feasible, effective, and interpretable way to classify gliomas and diagnose tumors in real time in clinical settings. MRI-Glioma Net represents a major step forward in neuro-oncology imaging; its strong performance across segmentation and grading tasks suggests it could be useful for pre-surgical planning, prognosis, and monitoring treatment response.
[...] Read more.Traumatic injuries in emergency response scenarios require rapid and reliable assessment methods capable of supporting medical prioritisation and efficient resource management. This paper presents a multispectral edge deep learning system designed for early injury detection, assistance prioritisation, and financial-resource decision support in emergency response environments. The proposed approach integrates RGB, thermal, near-infrared (NIR), depth, and contextual information within an edge-based analytical framework that combines multispectral data fusion, adaptive attention-based weighting of sensor modalities, convolutional neural networks for feature extraction and lesion segmentation, severity classification, triage index estimation, and resource allocation optimisation. The developed mathematical model describes the interaction among sensory processing, local inference, decision-making, and feedback-based model updating, and system stability is analysed using a Lyapunov-based approach under specified modelling assumptions. The system was evaluated on a simulated dataset containing N = 1000 multispectral observations, including 150 test samples for final performance assessment. The obtained results demonstrated promising performance on simulated data, achieving Accuracy = 0.912 ± 0.015, Recallcrit = 0.941 ± 0.021, Precisioncrit = 0.876 ± 0.024, and F1-score = 0.908 ± 0.018. The average edge inference delay was 425 ms, satisfying the considered real-time processing threshold of 500 ms. Compared with the baseline emergency response scenario without multispectral edge-based decision support, the proposed approach reduced triage decision errors by 18.6% and unproductive resource utilisation costs by 16.4%. The contribution of this research is the formulation of an integrated multispectral edge deep learning framework that combines injury detection, automated triage assessment, resource optimisation, and financial risk considerations within a unified decision-support architecture. The results indicate the potential of the proposed approach to improve emergency response processes. However, further validation on real-world clinical and operational datasets is required to confirm its practical applicability.
[...] Read more.Brain tumor refers to the abnormal enhancement of a collection of tissues or cells within the brain, occurring when brain cells grow improperly. Brain Tumors have currently been recognized as a major cause of mortality worldwide. As the most fatal category of cancer, brain tumors make detection and treatment critical for protecting lives. Single-Photon Emission Computed Tomography (SPECT) provides important imaging insights for detecting and evaluating brain cancer. However, variations in medical scenarios and imaging techniques make it difficult to improve models that are universally applicable across domains. This research suggests a novel framework for improving SPECT-based brain tumor diagnosis by combining Domain Adversarial Neural Networks (DANN) with Model Agnostic Meta Learning (MAML) and anisotropic diffusion filtering. The anisotropic diffusion filter reduces noise while maintaining tumor relevant structures, thereby preserving critical spatial properties for analysis. DANN supports domain adaptation by learning domain-invariant representations, allowing the model to perform consistently across different datasets. MAML enhances the model’s capability to adapt to new, unfamiliar domains with insufficient data, hence boosting its therapeutic value. Experimental results of the proposed DANN-MAML model with anisotropic diffusion filtering over diverse SPECT datasets outperformed existing strategies regarding both accuracy and flexibility, making it a promising tool for brain tumor identification.
[...] Read more.This paper presents Block-Based Compressive Sensing (BCS) of images using both traditional and learning-based approaches. An Image Block Mapping Lemma, formulated as a modified version of the Johnson-Lindenstrauss Lemma for image data, is proposed in this work, which justifies distance preservation between image blocks during mapping from the original domain to the compressed domain in block-based compressive sensing. The paper provides both theoretical proof and experimental validation of the proposed lemma. To quantitatively analyse distance preservation between image blocks in the original and compressed domains during block-based compressive sensing using traditional approaches, a new metric termed Distance Ratio (DR) is introduced. The difficulty of obtaining accurate real-time reconstruction using conventional block-based compressive sensing methods has encouraged the transition toward adaptive deep learning approaches. For learning-based sensing and reconstruction, the existing AutoBCS framework is enhanced by introducing an Initial Reconstruction Refinement Network (IRR-Net) between the initial and final reconstruction stages. Using the proposed model, improved reconstruction quality is achieved, with average PSNR gains of 0.69-1.62 dB and SSIM improvements of 0.01-0.02 for sampling rates of 0.30, 0.25, 0.10, and 0.04 across multiple benchmark datasets compared with the baseline architecture. The proposed model introduces an additional 0.05 million parameters and increases the computational cost by 3.64 GFLOPs due to the residual refinement module, while reducing the reconstruction time compared with the baseline AutoBCS framework. Experimental results demonstrate that the proposed model exhibits improved preservation of complex structures, edges, and fine details, particularly for images containing rich textures, dense structures, and high-frequency content.
[...] Read more.Adsorption of dyes present in textile industry waste water is an important process for environmental remediation and pollution control. Dyes are complicated organic compounds used in textile dyeing processes. The presence of dyes in wastewater is of major environmental concern because of their recalcitrance and potential toxicity. Adsorption technology is a promising solution for the removal of dyes from textile wastewater. It uses adsorbent materials to capture and immobilize dye molecules from aqueous solution. Machine learning techniques have been used to reduce the analytical efforts for prediction of concentration of dye solution in parts per million using a UV spectrophotometer.
In the research work, a prediction technique using Machine Learning and image processing is used to identify the color of images of samples and corresponding adsorption of the dye. The suggested model has achieved correct dye concentration estimates using retrieved image attributes with 93% accuracy with experimental data.
This paper introduces BlinkFusion, a real-time, interpretable, and detector-independent pipeline for analyzing blink events. The system finds eye ROIs using YOLOv5 as the main detector and a Haar cascade as a backup that has been calibrated. It then stabilizes the detections using lightweight tracking. A small landmark regressor inside each ROI gives six points to calculate the Eye Aspect Ratio (EAR), which keeps the geometric meaning. An uncertainty-weighted smoother combines pose, landmark, and detector confidence. An online hysteresis state machine with minimum-duration and derivative gates makes blink onsets and offsets. Platt-calibrated and fused detector confidences make strong arbitration possible at a reasonable cost. Performance is assessed on EyeBlinkDB (RGB ≥25 fps, ~720p), using subject-independent 10-fold splits and temporal-IoU matching. The model achieves Precision 0.928 ± 0.008, Recall 0.945 ± 0.008, F1 0.936 ± 0.008, AP@tIoU=0.3 = 0.962 ± 0.006, with timing precision of 28.1 ± 1.7 ms onset MAE and 37.1 ± 2.3 ms offset MAE. Condition stratification validates robustness: The values for Bright/Dim F1 are 0.947/0.922, for Frontal/±15°/≥30° yaw F1 they are 0.952/0.936/0.903, for Glasses (No/Yes) they are 0.946/0.925, and for Occlusion (No/Yes) they are 0.948/0.906. With adaptive scheduling (YOLO invocation rate ρ≈0.045), pipeline runs at about 32 FPS on the CPU (about 90 FPS on the observed frame loop) with a fixed 3-frame delay, which is fast enough for real-time use on cheap hardware. Ablations demonstrate progressive improvements resulting from calibrated fusion, tracking, quality-weighted EMA, and hysteresis (F1: 0.903 → 0.952). The approach is modular (you may switch the detector), doesn't need bounding boxes, and makes judgments that can be explained using EAR traces and thresholds. So, BlinkFusion is a useful, ready-to-use solution for HCI and clinical contexts that need clear, accurate, and quick blink analytics.
[...] Read more.