International Journal of Engineering and Manufacturing (IJEM)

IJEM Vol. 16, No. 5, Oct. 2026

Cover page and Table of Contents: PDF (size: 754KB)

Table Of Contents

REGULAR PAPERS

Comparative Evaluation of Fine-Tuned Transfer Learning CNN Architectures for Automated Citrus Fruit Disease Classification

By Nagineni Venkata Sireesha Gillala Rekha

DOI: https://doi.org/10.5815/ijem.2026.05.01, Pub. Date: 8 Oct. 2026

Early and accurate detection of citrus diseases is essential for maintaining fruit quality, minimising crop losses, and supporting sustainable agricultural production. Automated image-based diagnostic systems offer a scalable alternative to conventional manual inspection, which is often time-intensive, subjective, and susceptible to diagnostic variability. This study presents a transfer-learning-based deep learning framework for automated classification of citrus fruit diseases using a curated image collection derived from Food and Agriculture Organisation (FAO) resources. The dataset was organised into two classification groups: lemon and orange. The lemon dataset initially comprised 208 images across four classes: Healthy, Canker, Mold, and Scab, while the orange dataset contained 2,240 images across two classes. To address data scarcity and improve model generalisation, targeted data augmentation involving zooming (0.8–1.2), rotation (35°), and horizontal flipping was employed, increasing the datasets to 1,600 and 4,000 images for lemon and orange, respectively. Five ImageNet-pretrained convolutional neural network architectures—VGG16, ResNet50, InceptionV3, DenseNet121, and EfficientNetB0—were fine-tuned and evaluated using a stratified 70:30 training–testing protocol with five-fold cross-validation. Model performance was assessed using accuracy, precision, recall, and F1-score under a standardized experimental configuration. The results demonstrate that InceptionV3 achieved the highest classification accuracy of 90.0% on the four-class lemon dataset, while DenseNet121 obtained the best accuracy of 93.0% on the binary orange dataset. These findings indicate that appropriate transfer learning and targeted augmentation can substantially improve classification performance and model generalisation, particularly in limited-data agricultural imaging scenarios. The proposed benchmarking framework enables systematic, controlled comparisons of fine-tuned CNN architectures under consistent preprocessing, augmentation, training, and evaluation conditions. The resulting models can be integrated into real-time citrus disease diagnostic systems and deployed on resource-constrained platforms, including mobile and edge-computing devices, supporting scalable precision-agriculture applications.

[...] Read more.
Self-Evolving Digital Twins for Industry 5.0: Autonomous Model-Plant Mismatch Correction and Human Centric Cognitive Integration in Cyber Physical System

By Gaurav Jitendra Shahane Sudhir Madhav Patil Hrishikesh Kishor Jadhav

DOI: https://doi.org/10.5815/ijem.2026.05.02, Pub. Date: 8 Oct. 2026

Industry 5.0 does not extend Industry 4.0; rather, it demands something structurally different. Monitoring and prediction are no longer sufficient. Modern production environments require Digital Twin (DT) architectures that correct themselves autonomously and treat operator expertise as a formal computational input. Two gaps in the current literature block this transition. The first is the absence of autonomous Model-Plant Mismatch (MPM) correction. Every reviewed DT framework retains at least one human dependency, such as scheduled re-identification, manual labeling, or engineer approval before changes are committed. This holds across 67 publications covering manufacturing robotics, Computer Numerical Control (CNC) machining, predictive maintenance, and process control. The second gap is architectural. Every reviewed DT interface pushes information toward the operator, but none route operator knowledge back into the computational model as a structured input. Two integrated frameworks address these gaps. The Adaptive Self-Evolving Digital Twin (ASEDT) closes Gap 1 through the Real-Time Anomaly Detection Engine (RADE) anomaly detection, Online Parameter Estimation Module (OPEM) Bayesian parameter estimation, Structural Model Adaptation Layer (SMAL) meta-learning structural adaptation and Evolution Audit and Rollback Controller- (EARC) governed rollback validation. The Human Centric Cognitive Synchronization (HCCS) framework closes Gap 2 through Operator Intent Capture Layer (OICL) observation classification, Explainable State Presentation Layer (XSPL) expertise stratified presentation and Shared Mental Model Alignment Engine (SMMAE) continuous Mental Model Distance Metric (MMDM) computation. Both are coupled through the Triadic Feedback Architecture (TFA). ASEDT is validated via Software-in-the-Loop (SIL) testing on a Siemens NX MCD (Mechatronics Concept Designer) four corner punching machine at 100 Hz. Across 12 MPM scenarios it achieves a 62.6-94.1% Root Mean Square Error (RMSE) reduction, a 30 ms detection latency, parametric convergence within 4.7-11.9 minutes and zero manual interventions. HCCS empirical validation through a planned longitudinal study (N = 40, counterbalanced, with Tobii eye tracking) is described in Section 5.6.

[...] Read more.
Modeling and Manufacturing 3D Geometric Figures for STEM Training of Engineering Students: A Staged Methodology and Case Study

By Oleksandr Derevyanchuk Serhiy Balovsyak Nataliia Ridei Mickolay Dominikov Hanna Kravchenko Oleksii Pshenychnyi

DOI: https://doi.org/10.5815/ijem.2026.05.03, Pub. Date: 8 Oct. 2026

This article presents a methodology for modeling and manufacturing three-dimensional (3D) geometric figures. The need for developing the methodology is substantiated, and its significance for the STEM training of future specialists in engineering specialties is demonstrated. An advantage of the methodology is the clear structuring of all its stages, levels, and steps, which simplifies the process of implementing STEM projects. Structurally, the methodology includes a preparatory stage and four main stages. Stage 1 involves designing the object model, which includes collecting information about its characteristics and developing a mathematical model. Based on the mathematical model, the main parameters of the investigated object (part) are determined and calculated. At Stage 2, the computer-aided construction of a three-dimensional model of the object is performed, taking into account the parameters determined at the first stage. The development of drawings and the 3D model is carried out in an appropriate software environment, in particular, in the AutoCAD computer-aided design (CAD) system. Based on the calculated parameters, different types of three-dimensional models of the object are created: wireframe, surface, and solid. Stage 3 involves the implementation of additive manufacturing, that is, the manufacture of a physical prototype of the investigated object using 3D printing. For models of significant overall dimensions or complex geometric shapes that complicate printing, preliminary decomposition of the model into separate parts is performed. After printing, the individual components are joined, and the resulting physical prototype undergoes appropriate post-processing. Stage 4 involves the validation of the  manufactured physical 3D model by assessing its conformity with the digital model of the object. Comparison of the physical and software models can be carried out based on the analysis of their photographic images using digital image processing methods. In addition, a visual assessment of the prototype and measurements of its main geometric parameters are performed. An example of the implementation of a STEM project aimed at the creation, investigation, and physical reproduction of a three-dimensional model of a small stellated dodecahedron by means of 3D printing is presented. According to the developed methodology, at the first stage, a mathematical model of the dodecahedron was developed. At the second stage, based on this model, wireframe, surface (polygonal), and solid (volumetric) models were sequentially developed. An important advantage of the proposed methodology is the step-by-step construction of the wireframe model, which begins with the construction of simple geometric primitives. At the subsequent stages, individual primitives and their groups are sequentially combined into more complex structural elements. This approach simplifies the construction process and enables the step-by-step creation of complex three-dimensional models. At the third stage, the model was decomposed, prepared for printing, and its individual components were manufactured using an FDM 3D printer. After that, the printed parts were assembled and the resulting product was post-processed. As a result, a physical model of the small stellated dodecahedron was manufactured, whose geometric dimensions and surface quality meet the requirements and objectives of the STEM project. A physical model of the small stellated dodecahedron was fabricated, with a pentagon edge length of b = 25.3 mm and a circumscribed sphere diameter of DS = 127.2 mm. The maximum dimensional deviation of the dodecahedron was 0.3 mm.

[...] Read more.
MA-Helformer: A Sentiment-Aware Adaptive Framework for Multi-Horizon Crypto Forecasting

By Preeti Pandey Geeta Sharma Mukesh Kumar

DOI: https://doi.org/10.5815/ijem.2026.05.04, Pub. Date: 8 Oct. 2026

The cryptocurrency market is highly non-linear, non-stationary and driven by sentiment volatility, which poses fundamental challenges for price prediction. The recently proposed Helformer model couples a Transformer-based architecture with Holt-Winters exponential smoothing and shows excellent performance for univariate forecasting. However, the model relies only on closing price data and does not incorporate external sentiment sources, and uses static decomposition parameters without taking into account market conditions. Finally, the Helformer model has only been evaluated on a single asset. To systematically address these limitations, this paper proposes the Multivariate Adaptive Helformer (MA-Helformer), an improved deep learning framework with three main contributions. First, the input feature space is enriched with a complete multivariate set of 13 technical indicators, six on-chain metrics, seven macroeconomic variables, and three sentiment scores from FinBERT. Second, an adaptive market regime detection module based on Hidden Markov Models (HMM) dynamically conditions Holt-Winters decomposition parameters on detected market state (Bull, Bear, Sideways). Third, a multi-horizon forecasting head is trained jointly to produce price forecasts for 1-day, 7-day, and 30-day horizons simultaneously. We experiment on Bitcoin (BTC), Ethereum (ETH) and Solana (SOL) using daily data from January 2017 to December 2024. MA-Helformer reduces next-day RMSE by 49.2% relative to the baseline Helformer, achieves a Sharpe Ratio of 2.34 in a realistic trading simulation that accounts for variable transaction costs and slippage, and shows particularly strong improvements in Bear market regimes. All reported performance gains are statistically significant ("p<0.01" ) according to the Diebold-Mariano test. Experimental results demonstrate that MA-Helformer consistently outperforms the baseline Helformer and several competitive forecasting models across multiple evaluation metrics and forecast horizons, while achieving statistically significant improvements under the adopted experimental protocol. Although the proposed framework demonstrates promising predictive performance, its evaluation is based on historical market data and simulated trading conditions; therefore, further validation using real-time deployment and additional financial assets remains future work.

[...] Read more.
Grape Leaf Doctor: Severity Assessment of Grape Black Measles via DeepLabV3+ Segmentation and Interval Type-2 Fuzzy Logic

By Dipak Kumar Jana Sourav Mandal Sudipta Roy

DOI: https://doi.org/10.5815/ijem.2026.05.05, Pub. Date: 8 Oct. 2026

Automated diagnosis of Grape Black Measles (GBM) disease has become a pivotal aspect of modern agribusiness due to its efficiency and rapidity. Manual segmentation and diagnosis of GBM are intricate tasks due to time and cost constraints. In this study, we introduce Grape Leaf Doctor, a novel method for the automatic detection and severity analysis of GBM, utilizing the response surface approach and interval type 2 fuzzy logic inference (IT2FL). Firstly, we employ the DeepLabV3+ semantic segmentation model based on ResNet50 to perform pixel-level predictions on images of grape leaves affected by fungal lesions. This model enables the identification of “regions of interest” (ROIs) and the calculation of the percentage of infections (POI). Subsequently, the IT2FL rule-based system is constructed to assess the severity of disease damage based on these features. In the IT2FL system, Gaussian and trapezoidal “membership functions” (MFs) are explored for inputs and outputs to facilitate fuzzy inference and defuzzification. The severity of GBM infection is categorized into four levels: ‘Healthy’, ‘Mild’, ‘Medium’, and ‘Severe’. The experimental results on the IT2FL hold-out test dataset show a general classification accuracy of 98.34%, whereas RSM achieves 90.69%. By merging image processing and statistical modeling, the DeepLabV3+ framework of the IT2FL system can efficiently recognize GBM across varying disease risks.

[...] Read more.
Region-Specific Vehicle Classification in Bangladesh: A Comparative Study of CNN Architectures and Ensemble Strategies

By Shamsun Nahar Md. Nazim Hossain Noshin Un Noor Mohammad Anwar Hossain Md. Sabbir Hossain Babu Md. Sohorab Hossen

DOI: https://doi.org/10.5815/ijem.2026.05.06, Pub. Date: 8 Oct. 2026

The rapid growth of vehicles in Bangladesh has exacerbated traffic congestion, safety concerns, and administrative difficulties, creating an urgent need for Intelligent Transportation Systems (ITS). However, existing global datasets poorly represent Bangladeshi vehicles, and no region-specific detection models are currently available. To address this gap, this research constructs a balanced augmented dataset combining two public repositories, yielding images across ten vehicle classes. A harmonized dataset of manually verified images was split prior to augmentation, expanding the training set to 24,696 samples. Five deep Convolutional Neural Network architectures are then evaluated using transfer learning and extensive augmentation, and a convex ensemble of three selected models with optimized weights is developed to enhance robustness. The single ConvNeXt model achieves 98.980% ± 0.198% accuracy, while the ensemble attains 98.991% ± 0.156% accuracy. Statistical testing confirms the significance of these results; the ensemble effectively reduces misclassifications between visually similar categories. Overall, the proposed system provides a dependable platform for transportation applications in Bangladesh. Future work will address current limitations in static image and category scope through video-based analysis and lightweight deployment strategies.

[...] Read more.
Evolution and Clinical Translation of MRI Enhancement Methods: From Classical Filters to Deep Learning

By Idowu Olumayowa Ayodeji Amusa Kamoli Akinwale Ifeoluwa David Solomon Abolaji Okikiade Ilori

DOI: https://doi.org/10.5815/ijwmt.2026.05.07, Pub. Date: 8 Oct. 2026

Magnetic Resonance Imaging (MRI) is a fundamental diagnostic imaging modality that provides excellent soft-tissue contrast without ionizing radiation. However, MRI image quality is frequently degraded by noise, intensity non-uniformity (bias field), low spatial resolution, and motion artifacts, which adversely affect diagnostic accuracy and the performance of downstream artificial intelligence (AI) applications. This review presents a systematic and comparative assessment of MRI image enhancement techniques using a structured literature screening methodology inspired by the PRISMA framework. Unlike previous reviews that primarily summarize individual enhancement approaches, this study proposes a unified classification framework encompassing traditional image processing, deep learning (DL)-based, and hybrid enhancement techniques. A cross-paradigm comparison is performed using common evaluation dimensions, including enhancement accuracy, computational complexity, data requirements, generalization capability, interpretability, hardware dependency, and clinical readiness. The review further examines representative algorithms, benchmark datasets, quantitative performance metrics, clinical validation studies, regulatory pathways, and current challenges such as domain shift, explainability, and reproducibility. Emerging trends, including transformer-based architectures, self-supervised learning, multimodal enhancement, federated learning, and edge AI, are also discussed. The analysis indicates that although DL approaches consistently achieve superior quantitative performance, hybrid techniques provide a more balanced trade-off between enhancement accuracy, interpretability, computational efficiency, and deployment feasibility. This review offers a comprehensive reference by integrating methodological, technical, and clinical perspectives while identifying key research directions for developing robust, trustworthy, and clinically deployable MRI image enhancement systems.

[...] Read more.
IIoT-RoboAuth: PUF-Based Command-Bound Authentication for Industrial Robotic Control Networks

By Haewon Byeon

DOI: https://doi.org/10.5815/ijem.2026.05.08, Pub. Date: 8 Oct. 2026

Industrial robotic controllers require identity assurance and command-level integrity within bounded control intervals. This study redesigns a fog-assisted three-factor scheme as IIoT-RoboAuth, a PUF-rooted protocol in which the session key is derived from the three participant nonces, an ephemeral P-256 secret, the policy epoch, and the current RPUF epoch; each command is then authenticated over its canonical payload, sequence, timestamp, role scope, and safety-profile hash. The protocol uses four handshake messages (two end-to-end round trips), 483 application-layer bytes, and a 224-byte command envelope. Evaluation used a Python 3.12 message-driven simulator rather than robot hardware or NS-3. Thirty fixed seeds generated 1,000 sessions and 1,000 trials for each of four attack classes per seed. The modeled mean authentication latency was 7.903 ms (95% CI, 7.896-7.909 ms), compared with 22.735 ms (95% CI, 22.698-22.773 ms) for the centralized baseline under the stated delay assumptions. The complete validator rejected 30,000 of 30,000 command modifications, replays, stale commands, and out-of-range commands in each class; removing the command MAC, sequence chain, 5-ms freshness check, or kinematic check caused the corresponding attack class to pass. The result establishes reproducible protocol-level command binding and an explicit deployment boundary; hardware timing, PUF reliability, physical tamper resistance, and safety certification remain outside the evidence provided here.

[...] Read more.
Attention-Guided Deep Learning Framework for Ovarian Cancer Subtype Classification

By Vijay H. Kalmani Nagaraj V. Dharwadkar Amol C. Adamuthe Altaf Husain

DOI: https://doi.org/10.5815/ijem.2026.05.09, Pub. Date: 8 Oct. 2026

Ovarian cancer histotype classification is challenging because of substantial morphological heterogeneity and subtle subtype-specific features. This work presents a controlled evaluation of high-resolution pathology representations and slide-level aggregation methods for automated classification of five ovarian cancer subtypes from whole-slide histopathology images. Precomputed CONCH patch embeddings were aggregated using mean pooling, max pooling, gated attention-based multiple-instance learning, and a max-pooling cascade. The models were evaluated on 513 non-TMA whole-slide images from UBC-OCEAN using leakage-controlled five-fold cross-validation, with each slide receiving exactly one out-of-fold prediction. Gated ABMIL achieved a balanced accuracy of 81.39% and a macro-F1 score of 81.96%. Max pooling produced the highest numerical performance, with a balanced accuracy of 81.44%, macro-F1 of 82.59% (95% CI: 78.84-86.26%), macro-AUROC of 96.99%, and macro-AUPRC of 90.79%. However, paired slide-level bootstrap comparisons found no statistically significant differences among the CONCH aggregation strategies after Holm correction. Compared with the EfficientNet-B0 thumbnail baseline, CONCH max pooling improved macro-F1 by 30.30 percentage points and balanced accuracy by 25.70 percentage points, with both paired bootstrap confidence intervals excluding zero. Attention weights enabled visualization of influential patches, although these regions were not independently validated by pathologists. The findings show that high-resolution pathology foundation-model representations support ovarian cancer subtyping, while greater aggregation complexity does not necessarily improve performance. External multi-institutional validation is required before clinical generalizability can be established.

[...] Read more.
Leakage-Controlled Recording-Level Evaluation of Machine Learning, Deep Learning, and Convolutional-Kernel Models for EEG-Based Epileptic Seizure Detection

By Sanagavarapu Sunitha Umadevi Ramamoorthy

DOI: https://doi.org/10.5815/ijem.2026.05.10, Pub. Date: 8 Oct. 2026

Electroencephalogram (EEG)-based seizure detection is a challenging time-series classification problem because EEG signals are nonlinear, non-stationary, and temporally heterogeneous. This study presents a unified comparison of classical machine-learning, neural and deep-learning, convolutional-kernel time-series, and probability-level ensemble models using the Epileptic Seizure Recognition benchmark dataset. The 11,500 EEG segments were reconstructed into 500 recording groups, and recording-grouped nested cross-validation was applied to prevent segments from the same recording from crossing training and test partitions. Depending on the modelling paradigm, segments were represented either by 67 handcrafted time-, frequency-, and nonlinear-domain features or by ordered 178-sample EEG sequences. Recording-level average precision was used as the primary evaluation metric, supported by ROC-AUC, threshold-dependent measures, Brier score, bootstrap confidence intervals, ablation analysis, explainability, and computational-efficiency benchmarking. RBF-SVM achieved the highest average precision of 0.997985 and ROC-AUC of 0.999475. MultiRocket produced the lowest family-representative Brier score of 0.006148, while MLP achieved the lowest standalone inference latency after feature extraction. The selected meta-ensemble produced strong threshold-dependent performance but did not improve overall precision–recall ranking. Paired bootstrap analysis found no statistically significant differences among the family representatives after multiple-comparison correction. These findings demonstrate that handcrafted-feature and time-series models can both achieve strong recording-level performance when evaluated under a consistent, leakage-controlled protocol. Because verified patient identifiers were unavailable, the results should not be interpreted as evidence of patient-independent clinical generalization.

[...] Read more.
A Low Complexity Hybrid LPC-Wavelet-Golomb Audio Compression Framework for Real Time FPGA Implementation

By Jaimy James Poovely Raveena Judie Dolly

DOI: https://doi.org/10.5815/ijem.2026.05.11, Pub. Date: 8 Oct. 2026

Efficient audio compression with low computational complexity is essential for real-time embedded systems operating under tough latency, memory, and hardware resource constraints. This paper presents a low-complexity hybrid audio compression framework that integrates Linear Predictive Coding (LPC), discrete wavelet transform (DWT) and Golomb entropy coding into a unified pipeline-oriented architecture for real-time FPGA implementation. The proposed framework uses LPC for short-term spectral modeling and residual extraction, Daubechies 4 wavelet transform for multi-resolution energy compaction and Golomb entropy coding for efficient compression of the resulting coefficients. The entropy coding scheme is lightweight and agreeable to hardware implementation. The architecture uses fixed-point arithmetic and pipelined processing to provide deterministic execution with low computational complexity on an Artix-7 FPGA platform. Experimental evaluation was performed on a 16 kHz uncompressed speech signal with 20 ms frames. The proposed framework achieved a compression ratio of 5.51 which is higher than that of LPC only (2.00), wavelet only (2.00) and MP3 (5.33) under the same evaluation conditions. The hardware implementation only used 3.83% LUT utilization, 0.94% flip-flop utilization and one DSP block. The processing latency of 0.037 ms per frame is significantly less than the 20 ms frame duration for real-time operation. Objective evaluation yielded SNR of 69.37 dB, STOI of 0.990, and PESQ of 2.19. This demonstrates that the suggested framework favors compression efficiency and hardware simplicity at the cost of reasonable reconstruction quality. The proposed LPC-Wavelet-Golomb architecture provides a practical compromise between compression performance, implementation complexity and real-time FPGA suitability for embedded audio compression applications.

[...] Read more.
The Efficiency of FPGA-based Matrix Multiplication Accelerators for Neural Network Algorithms

By Yaroslav Klyatchenko Oxana Tarasenko Klyatchenko

DOI: https://doi.org/10.5815/ijem.2026.05.12, Pub. Date: 8 Oct. 2026

The contemporary challenges facing computer engineering have led to a focus on improving the efficiency of hardware and computing systems, as well as on optimising the operation of devices for effective data processing. In most modern systems of artificial intelligence, computer vision and digital signal processing, matrix multiplication is a basic operation. With the increasing resolution of sensors and the growing complexity of neural networks, classical general-purpose processors face the ‘von Neumann bottleneck’, where the data transfer rate between memory and the processor is limited, and the sequential execution of instructions does not allow the required real-time throughput to be achieved. The subject of this research is a System-on-a-Chip architecture that combines a dual-core processor based on the ARM architecture with programmable logic. This hybrid structure allows the most computationally intensive tasks to be offloaded to the hardware, whilst leaving control and the implementation of high-level interfaces to software. A neural network accelerator based on programmable logic devices offers advantages such as the capability for stream processing, which minimises the number of accesses to external memory, and the ability to utilise so-called mixed-precision computing. Furthermore, the accelerator’s efficiency is achieved through the use of multiple processing elements, enabling parallel computation of the neural network’s output channels. It has been demonstrated that implementing matrix operations on programmable logic enables parallelism of hundreds of operations per clock cycle, which is unattainable for systems based on general-purpose CPUs of the same class. The proposed implementations for a programmable logic-based accelerator have demonstrated speed-up compared to a system based on an ARM-architecture processor, as well as high adaptability to Deep Learning tasks.

[...] Read more.
Threshold-Based Active Cooling and LDR-Driven Dual-Axis Tracking for Low-Cost PV Performance Enhancement: Hardware Implementation and Experimental Evaluation

By SriLakshmi Lavanya Kota M. Raja Nayak M. Sudheer Kumar T. Vamsee Kiran Pradeep Panthagani B. Devulal Harish Sesham

DOI: https://doi.org/10.5815/ijem.2026.05.13, Pub. Date: 8 Oct. 2026

Solar energy is a widely utilized renewable source, yet the performance of photovoltaic (PV) systems is significantly affected by temperature rise, dust accumulation, and improper panel orientation. While many approaches, such as temperature management, water cooling, and active tracking, whether used alone or in combination, have been adopted to improve PV cell efficiency, they have provided only limited enhancement. This study presents a low-cost solar PV module performance-enhancement prototype incorporating threshold-based active cooling, LDR-driven dual-axis solar tracking, automated surface cleaning, and monitoring functionality. The proposed system employs a reliable ATMEGA328 controller well suited for hybrid intelligent function of regulation of DC fan and water-cooling mechanisms based on real-time temperature data, ensuring activation when the panel temperature exceeds the limit of 35°C. Dual-axis tracking using LDR sensors and servo motors optimizes solar irradiance absorption, while IoT intelligence connectivity enables remote monitoring, data visualization, and system diagnostics. Additionally, Bluetooth support operation aids in manual operation control of the cooling system. The performance was evaluated against a conventional fixed PV configuration using the recorded daytime power-output profile. The proposed integrated configuration increased the cumulative measured power output by approximately 18.8% relative to the conventional reference, with the instantaneous enhancement reaching approximately 36% during the evaluated high-output operating period; the prototype temperature measurements ranged from approximately 35 to 44 °C under the recorded test conditions. The results demonstrate the potential of coordinated solar tracking and thermal management to improve PV power generation.

[...] Read more.
Multi-Resolution Attention-Fused, Calibration-Aware Diagnosis of Kidney CT (Normal/Cyst/Stone/Tumour)

By Shanker M. C. N. Sankar Ram V. Gokula Krishnan Pinagadi Venkateswararao Therasa Michael S. Kaviarasan

DOI: https://doi.org/10.5815/ijem.2026.05.14, Pub. Date: 8 Oct. 2026

Accurate classification of kidney abnormalities from computed tomography (CT) images is essential for early diagnosis and clinical decision-making. However, distinguishing kidney stones, cysts, tumours, and normal kidneys remains challenging because of variations in lesion size, appearance, and imaging conditions. This study proposes a multi-resolution attention-fused, calibration-aware deep learning framework for four-class kidney CT image classification. The framework integrates convolutional neural networks and a Vision Transformer to extract complementary local and global features, which are combined through an adaptive attention-based fusion mechanism. Calibration-aware learning, incorporating focal loss, multi-class Brier loss, and feature-level orthogonality regularization, is employed to improve prediction reliability and confidence estimation. The model was trained and evaluated using an image-level data split on a publicly available kidney CT dataset containing normal, cyst, stone, and tumour images. Experimental results demonstrate an overall classification accuracy of 99.3%, a Macro-F1 score of 99.1%, an AUROC of 0.999, and an AUPRC of 0.998 on the independent test set. Robustness analysis under synthetic image corruptions showed only a modest reduction in performance, with the AUROC remaining at 0.995 under severe motion blur conditions. External validation on two independent cohorts (External Site-A and External Site-B) achieved AUROC values of 0.993 and 0.988, respectively, while post-hoc temperature scaling improved calibration by reducing the Expected Calibration Error (ECE) from 0.028 to 0.012 and 0.034 to 0.015. These results demonstrate that the proposed framework provides accurate, robust, and well-calibrated predictions, supporting its potential application in computer-aided kidney CT image analysis.

[...] Read more.
Adaptive Myoelectric Prosthetic Control Using Hybrid CNN–Vision Transformer–LSTM Networks and Federated Reinforcement Learning

By P. S. Saritha S. Poonguzhali B. Mohan

DOI: https://doi.org/10.5815/ijem.2026.05.15, Pub. Date: 8 Oct. 2026

EMG-based prosthetic control systems are known to suffer from muscle fatigue, electrode displacement, signal drift, and inter-subject variability issues, compromising their robustness and long-term stability. While recently developed deep learning techniques deliver highly competitive gesture recognition accuracy, most existing solutions lack adaptive learning and privacy-preserving model optimization mechanisms required for the real-world implementation. In this paper, we propose a new EMG-based adaptive prosthetic control system combining the Hybrid CNN, ViT, and LSTM architectures and applying federated reinforcement learning (FRL). The CNN component of the proposed architecture extracts local spatial information from multi-channel EMG signals, the Vision Transformer analyzes inter-channel dependencies using a self-attention mechanism, while the LSTM models the temporal dynamics of muscle activation. Additionally, the reinforcement learning algorithm constantly updates control policies according to user feedback, while federated learning allows users to collaborate on model optimization without exposing raw EMG data, thus protecting user privacy. In order to enhance the system robustness in real-world conditions, the Adaptive Muscle Signal Learning (AMSL) technique is used to counteract the adverse effect of signal drift, muscle fatigue, and electrode displacement. An EMG dataset specification, called EMGPro-2026, is introduced to establish the criteria of data acquisition, gesture recognition classes, preprocessing, and evaluation process. Based on computational assessment and simulation analysis, there is proof of concept for the proposed framework that enables high gesture classification accuracy, adaptation for personalization, and efficient learning with privacy preservation for real-time use in prosthetics. The proposed CNN-ViT-LSTM with federated reinforcement learning presents an intelligent solution for the future generation of myoelectric prosthetic control systems. Our future work will entail the validation of our framework using EMG signals recorded ethically. Because the current research provides a conceptual framework, there has been no experimental validation of the performance using human EMG datasets. However, based on the benchmarks developed from the existing state-of-the-art literature, it can be said that the suggested framework for Hybrid CNN–ViT–LSTM with Federated Reinforcement Learning is predicted to deliver approximately 94.5% gesture classification accuracy, which will be better than the conventional techniques such as CNN, CNN–LSTM, and SVM and will also increase adaptability, personalization, and privacy preservation.

[...] Read more.
Field Performance and Degradation Comparison of ESS Multicrystalline and Conventional Polysilicon PV Modules in Hot Indian Climates: A One-Year Rooftop Study

By Ramchander Nirudi G. Tulasi Ram Das T. S. Surendra

DOI: https://doi.org/10.5815/ijem.2026.05.16, Pub. Date: 8 Oct. 2026

Photovoltaic (PV) technology is crucial for sustainable energy generation, but the performance of PV systems in real-life conditions is closely associated with the material quality, degradation and the climatic conditions. While conventional poly-Si modules are already widely installed, Elkem Solar Silicon (ESS®)-based multicrystalline silicon modules are more sustainable owing to energy savings in silicon production. But little field data is available on their behaviour in the hot Indian climate. Thus, the field performance of ESS® multicrystalline silicon and conventional poly-Si modules after one year of installed operation in a rooftop grid-connected 6.71 kWp PV system at BVRIT, Telangana, India is assessed. The demonstration system comprises 28 modules (organized in four rows of seven, with 14 ESS® modules and 14 conventional poly-Si modules). Performance analysis was carried out using field-based energy yield, I–V and P–V characteristics, peak power degradation, electrical mismatch, as well as electroluminescence imaging of the modules. The findings show that both technologies had minimal degradation, with an average annual power loss of 0.3% for ESS® modules and 0.4% for conventional poly-Si modules. ESS® module technology showcased 1-1.5% higher energy yield than conventional poly-Si modules, mainly because of better high-temperature operation. At the module level, it was detected that the majority of the changes in power were due to variations in current (Isc and Imp) rather than in Voc = 37.56 V and Vmpp = 30.23 V. Electroluminescence image analysis showed only minor defects from handling, but no defects caused by material. The results confirm that ESS® modules offer higher energy yield and are more suitable for sustainable grid-connected PV systems.

[...] Read more.
Spatially Adaptive Infrared–visible Fusion with EfficientDet for Night-time Multi-class Intrusion Detection

By Rithick S. Jenefa A. Abirami M. K. G. Sheeba Merlin Antony Taurshia Lincy A.

DOI: https://doi.org/10.5815/ijem.2026.05.17, Pub. Date: 8 Oct. 2026

Night-time perimeter monitoring in low illumination remains challenging because cluttered terrain, partial occlusion, and thermal crossover conditions simultaneously distort object boundaries and visual cues. Visible spectrum sensing loses discriminative texture at night, while infrared sensing preserves thermal salience but lacks structural context. Conventional pipelines that rely on single-modality detection or simple blending commonly produce unstable recall or avoidable false alarms when the background changes. The present paper presents an EfficientDet Fusion Intrusion Detector (EFID) which combines registered infrared and visible frames using an attention-guided weighting module and a spatially adaptive activity-weighted pixel fusion step. The feature-level attention weights estimate the reliability of each modality and guide the activity-weighted pixel fusion stage. The fused RGB image is resized from 1024 × 768 to 896 × 896 before being processed by EfficientDet-D3. The resultant fusion retains visible structural edges, while infrared target evidence is simultaneously enhanced. The resulting three-channel representation is processed by an EfficientDet D3 detector augmented with a bidirectional feature pyramid network (BiFPN) for multi-scale localisation under night conditions and clutter. Experiments use the public Multi-scenario Multi-modality Fusion and Detection (M³FD) benchmark containing 4,200 aligned infrared and visible pairs at 1024×768 resolution, annotated with 33,603 bounding boxes across six classes. The proposed EFID achieves 91.7% mAP@0.5 and 67.3% mAP@0.5:0.95, with 93.4% precision and 89.8% recall at 34.2 FPS on an NVIDIA RTX 3080. Systematic ablation confirms the independent contribution of the attention module, the activity-weighted pixel fusion term, and each BiFPN iteration. Cross-dataset experiments demonstrate that zero-shot transfer retains 84.1% mAP@0.5, rising to 86.0% with 10% target fine-tuning. The results indicate practical night robustness and real-time feasibility, while extreme occlusion and calibration drift remain limiting factors that motivate alignment-aware training and lightweight deployment optimisation as future work.

[...] Read more.
CPU-Centric Evaluation of Lightweight CNNs for Four-Class Chest X-Ray Classification: Performance, Latency, and Quantization Limits

By Kateryna P. Hazdiuk Roman V. Movcheniuk

DOI: https://doi.org/10.5815/ijem.2026.05.18, Pub. Date: 8 Oct. 2026

In most studies applying deep learning to medical image analysis, classification accuracy is the sole optimization criterion, while the computational cost of the resulting models remains undocumented: parameter and operation counts are reported only occasionally, and inference latency on a central processing unit is almost never published. This makes an informed model choice impossible for institutions without graphics accelerators, such as district hospitals, mobile diagnostic units and field hospitals. This paper experimentally investigates the trade-off between diagnostic performance and computational resources for four-class chest X-ray classification. The computational cost of 15 widely used architectures was evaluated with a single tool at an input resolution of 224×224, and the gap between the heaviest and the lightest architecture reaches a factor of 323. Four models – a heavy baseline, two lightweight architectures, and a custom compact network – were trained under a unified protocol on the COVID-19 Radiography Database (21,165 images across four classes) and profiled strictly on a CPU with a batch size of one. The three standard architectures were initialized with ImageNet-pretrained weights, and the custom network was trained from scratch. Each architecture was trained across three independent random seeds on a fixed data split, so all performance metrics are reported as mean ± std, and every comparison is accompanied by a confidence interval and a significance test. The lightweight architecture trailed the baseline by only 1.24 percentage points in macro-F1 (95% CI: 0.84 to 1.63, p = 0.005) while requiring 69 times fewer operations, achieving 19.3 times lower latency, and reducing model size by a factor of 15.4; the efficiency metric, defined as macro-F1 per GFLOP, differs by a factor of 68. Additionally, post-training quantization systematically failed across all three runs for both architectures combining depthwise convolutions with Squeeze-and-Excitation blocks, whereas the same pipeline quantized the remaining models without statistically significant loss. Neither depthwise convolutions nor the Hard-Swish activation accounted for the failure; instead, the number of channel-wise multiplication nodes in the exported graph separated the two groups exactly, making the risk identifiable directly from the static graph prior to training. The practical implication is that the quantizability of a lightweight model must not be assumed; it must be verified empirically.

[...] Read more.
From Sentiment to Cognition: Task-Driven Deci-sion Framework for Transfer Learning in 4-Class Cognitive Thought Classification

By Jitendra Singh Geeta Sharma

DOI: https://doi.org/10.5815/ijem.2026.05.19, Pub. Date: 8 Oct. 2026

Traditional sentiment analysis performs well in terms of polarity detection, but it cannot capture cognitive nuances needed for mental health monitoring. In this paper, we address two gaps in the current literature: the lack of task-driven guidance on how to choose transfer learning strategies and the lack of validation on cognitively rich da-tasets, other than standard polarity benchmarks. We present a systematic review of 150 transfer learning studies follow-ing PRISMA guidelines, summarising the findings into a decision framework that maps data availability, domain spec-ificity, and resource constraints to optimal model selection. We empirically validate our decision framework on a novel 4-class cognitive thought classification dataset with human-elicited samples and GPT-4 augmented data. Active learn-ing reduced annotation effort by 60% and achieved high inter-rater agreement (Fleiss’ Kappa = 0.82). According to our framework, 12 models were evaluated. BiLSTM-TF-IDF achieved 93.9% macro F1, outperforming transformer models by 0.5% with 9× less compute, supporting the medium-data/domain-specific path. Specifically, on our 4,944-sample domain-specific cognitive thought dataset using a single NVIDIA RTX 3090 GPU, BiLSTM-TF-IDF outperformed RoBERTa-base by 0.5% macro F1 (93.9% vs. 93.4%, McNemar p=0.004) via uncertainty-sampling active learning with SVM confidence threshold <0.40, while requiring approximately 9.4× less inference time and a 32× smaller model footprint (with a 6× lower peak GPU memory requirement). Uncertainty was quantified using least-confidence sam-pling (selecting samples where max class probability <0.40); the cross-entropy loss function was used for both the SVM-guided active learning and BiLSTM training. We also evaluate cross-dataset performance on IMDb and SST-2, demonstrating reasonable generalisability. Demographic analysis shows cognitive patterns that have implications for mental health. The dataset and code have been published at Zenodo (DOI: 10.5281/zenodo.17444289) to facilitate reproducible cognitive NLP research.

[...] Read more.
Predicting Social Media Addiction Levels Among University Students Using Machine Learning and SHAP Explainability

By Md Saiful Islam Md Safanur Islam Md. Rabbi Khan

DOI: https://doi.org/10.5815/ijem.2026.05.20, Pub. Date: 8 Oct. 2026

Overindulgence in social media use is increasingly becoming a matter of concern amongst students and is normally accompanied by behavioural, academic and health implications. The purpose of this study is to measure the rates of social media addiction among students and define the most important factors influencing it by means of data-driven methods. The analysis was done with a dataset of 705 student responses to determine the trends in addiction and forecast the level of addiction using machine learning models. The descriptive analysis indicated that 57.87 percent of students were in the High addiction category with the results being equal among all genders. Correlation analysis revealed that the greater the addiction scores, the greater the negative academic impact and the interpersonal conflicts, and conversely, the greater the influence on the sleep duration and mental health. The classification using several machine learning models was done, with K-Nearest Neighbours (KNN) and CatBoost having the greatest accuracy of 98.6 percent. To assess model robustness, the optimized KNN model was further validated using ten-fold cross-validation, achieving a mean accuracy of 97.17% (95% CI: 95.52%–98.81%). The optimized KNN model demonstrated stable predictive performance across different data partitions. SHAP-based explainability identified average daily social media usage, mental health score, sleep duration and conflicts over social media as the most influential predictors of social media addiction. These results indicate the necessity of awareness and careful intervention approaches and show the potential of machine learning to avert early social media addiction in students.

[...] Read more.
A Lightweight Face Anti-Spoofing Framework with Spoof Artifact Enhancement and Adaptive Feature Fusion

By Mudunuru Suneel Banothu Yedukondala Venkata Naga Raja Swamy Madhava Rao Maganti Seva Sreedhar Babu P. Rama Koteswara Rao Vijaya Kumari Devarapalli Kama Ramudu

DOI: https://doi.org/10.5815/ijem.2026.05.21, Pub. Date: 8 Oct. 2026

Face recognition is widely used for biometric authentication in applications such as mobile devices, financial services, intelligent surveillance, access control, and border security. However, face recognition systems remain vulnerable to presentation attacks, including printed photographs, replay attacks, and three-dimensional masks. Although recent deep learning-based face anti-spoofing (FAS) methods have achieved substantial improvements, many existing approaches still involve considerable computational cost, provide limited emphasis on fine-grained spoof artifacts, and face challenges in effectively integrating heterogeneous representations. To address these limitations, this paper proposes  
a Lightweight Face Anti-Spoofing Framework with Synthetic Multi-Modal Representation, Spoof Artifact Enhancement, and Adaptive Feature Fusion. The framework starts from a single RGB facial image and constructs complementary Depth and Near-Infrared (NIR) representations using a Depth and Near-Infrared Construction Module (DNCM), rather than requiring dedicated Depth or NIR sensors. A shared EfficientNetV2 backbone is then employed to extract features from the three representations with reduced computational redundancy. The proposed Spoof Artifact Enhancement Module (SAEM) emphasizes subtle spoof-specific visual cues, while Cross-Modal Consistency Learning (CMCL) reduces representation discrepancies across the constructed modalities. Subsequently, the Adaptive Feature Fusion Module (AFFM) dynamically weights the refined representations according to their discriminative contribution. Extensive experiments on CelebA-Spoof, CASIA-SURF, HQ-WMCA, and SiW-M demonstrate the effectiveness of the proposed framework. The framework achieves accuracies of 98.16%, 97.54%, 99.08%, and 97.18%, with corresponding ACER values of 2.87%, 4.36%, 2.03%, and 4.18%, respectively. Across five independent runs, statistical analysis further indicates consistent performance with significant improvements over the selected baseline. In addition, the framework requires only 9.8 million parameters and 2.1 GFLOPs and achieves an inference speed of 82 FPS, demonstrating a favourable balance between detection effectiveness and computational efficiency. The results indicate that synthetic multi-modal representation combined with explicit spoof artifact enhancement and adaptive feature fusion can provide an efficient solution for robust and real-time face anti-spoofing.

[...] Read more.
IBSMFO-Optimized UPQC-DG Control for Harmonic Mitigation in Hybrid Renewable Microgrids

By Shravani Chapala

DOI: https://doi.org/10.5815/ijem.2026.05.22, Pub. Date: 8 Oct. 2026

Distributed generation (DG) systems are increasingly integrated into modern power networks through power-electronic interfaces. However, the presence of nonlinear loads and power-electronic converters introduces harmonic distortion, which can significantly degrade power quality, particularly in three-wire systems. This study investigates harmonic mitigation in a hybrid renewable DG system incorporating solar, wind, fuel-cell-based generation, and grid-connected operation using a Unified Power Quality Conditioner (UPQC). An Improved Bat Search–Moth Flame Optimization (IBSMFO) based control strategy is proposed for optimizing the proportional-integral (PI) controller parameters of the UPQC. The IBSMFO algorithm combines the search characteristics of the Improved Bat Search Algorithm (IBSA) with the Moth Flame Optimization Algorithm (MFOA) to minimize the selected error-based objective function and enhance the dynamic performance of the compensation system. The proposed UPQC is evaluated under ideal and non-ideal source-voltage conditions with nonlinear loads. Simulation results demonstrate effective mitigation of voltage and current harmonics and improved power quality. In particular, the source-current THD is reduced to 4.03% under the operating conditions considered, while a minimum THD of 1.01% is achieved under the corresponding compensation condition. The results confirm that the proposed IBSMFO-optimized UPQC provides effective harmonic compensation and improved power quality performance in hybrid renewable DG systems.

[...] Read more.
A TD3-based Deep Reinforcement Learning algorithm for Voltage Regulation of Dual Active Bridge (DAB) DC-DC Converter under Triple Phase-Shift Modulation

By K. Girinath Babu J. Sivavara Prasad V. Vasudevan

DOI: https://doi.org/10.5815/ijem.2026.05.23, Pub. Date: 8 Oct. 2026

Owing to the bidirectional power transfer, galvanic isolation and high conversion efficiency, the Dual Active Bridge (DAB) converter has become a popular topology for bidirectional DC-DC power conversion. It applies to electric vehicles, battery energy storage systems and DC microgrids due to its properties. The non-linear input-output characteristics and parameter variations with operating conditions create a difficulty in controlling the output for accurate voltage regulation. A Deep Reinforcement Learning (DRL) control method is proposed to overcome the nonlinearity and voltage regulation issue of the DAB converter. Initially, a state space model and generalized averaging model are used to accurately model the converter dynamics. The control performance is explored in multiple phase-shift modulation techniques like Single Phase Shift (SPS), Extended Phase Shift (EPS), Dual Phase Shift (DPS) and Triple Phase Shift (TPS). A DDPG agent is first implemented for continuous phase-shift control; then, to improve the learning stability, convergence speed and voltage regulation performance, a TD3 agent is introduced. The performance of the proposed DRL is tested and compared to the Grey Wolf Optimizer tuned Proportional–Integral (GWO-PI) control and Model Predictive Control (MPC) for various input voltage and loading conditions. The simulation results show that the settling time of the proposed TD3 controller is 4.3ms, which is 35% lower than that of the MPC and GWO-PI controllers. The steady state voltage ripple of the proposed TD3 controller is reduced by 35% compared to the MPC and GWO-PI controllers. The voltage regulation accuracy of the proposed TD3 controller is better than that of the MPC and GWO-PI controllers.

[...] Read more.
A Systematic Review and Interaction Matrix of WEDM Process Parameters: Frequency Analysis and Research Gaps in Surface Integrity and Productivity

By Manojkumar Subrao Kate Priyaranjan Samal Kotthapalli Karthik

DOI: https://doi.org/10.5815/ijem.2026.05.24, Pub. Date: 8 Oct. 2026

Wire electrical discharge machining (WEDM) is a heat-based, non-contacting and precision machining process which can produce complex geometry in any conductive material regardless of its hardness. Although many studies on the investigations of individual input factors in the WEDM process have been done, no complete study on the frequency weighted correlation of all the key input factors with all the output parameters has been made yet. This paper tries to fill this gap through reviewing peer reviewed journal articles available on Scopus, Web of Science, Science Direct, Springer Link, Taylor & Francis Online, and IEEE Xplore according to PRISMA 2020 methodology. The number of key input and output factors is thirteen and eight respectively and they have been addressed in this paper. Frequency weighted analysis proves that Ton has been mentioned in 63%, Toff in 62%, V in 41%, Ip in 38%, Ra in 62%, and MRR in 48% of all papers. On the other hand, fatigue resistance, residual stresses, geometry deformation, and cutting speed still lack sufficient consideration even though they play a vital role in industry. The literature review also highlights powder mixed dielectrics, coated wire electrode, multi-objective optimization approach, artificial intelligence (AI), non-dominated sorting genetic algorithm (II), technique for order preference by similarity to ideal solution (TOPSIS), and machine learning as areas of emerging interest in improving the machining process. In general, the literature review serves as a good source for understanding the important process variables, current research trends, and gaps, and can be used for multi-objective process optimization and future research.

[...] Read more.
Attention-Based Spatial-Temporal Learning for Enhanced Kidney Abnormality Diagnosis from Ultrasound Images

By Ravikumar Ch. Satyanarayana Nimmala P. Vasanthsena Nenavath Chander R. Sahith

DOI: https://doi.org/10.5815/ijem.2026.05.25, Pub. Date: 8 Oct. 2026

Kidney abnormalities, when detected late, can develop into chronic conditions such as kidney tumors and cardiorenal syndrome. Recently, there has been a move toward automation, especially as it pertains to developing a framework for advanced preprocessing, feature fusion, and hybrid spatial-temporal learning for the detection of the abnormalities in question. For example, Non-Local Means (NLM) filtering for noise reduction, image quality improvements through U-Net kidney segmentation and deep learning-based text annotation removal, as well as robust training through data augmentation and preprocessing ends up facilitating the process. The combination of classical texture and shape descriptors and deep embeddings from a Vision Transformer (ViT) through an attention mechanism allows the model to pinpoint where clinically relevant focus should be. A hybrid CNN-LSTM framework attends to spatial and temporal feature extraction, while attention modules receive the classification output to refine it. The incorporation of weighted binary cross-entropy for handling class imbalance, and of explainable AI, is demonstrated through Integrated Gradients. Experimental results for VATLA show it used advanced algorithms and automation principles in developing systems for all previous performance indicators, while providing evidence of scoring 0.94 in accuracy, 0.92 in precision, 0.95 in recall, and 0.96 in AUC, or area under the curve.

[...] Read more.