International Journal of Engineering and Manufacturing (IJEM)

IJEM Vol. 16, No. 4, Aug. 2026

Cover page and Table of Contents: PDF (size: 751KB)

Table Of Contents

REGULAR PAPERS

Hybrid Quantum-Classical Framework for Computational Mental Energy from Multichannel EEG Streams

By Mykhailo Vernik Liubov Oleshchenko

DOI: https://doi.org/10.5815/ijem.2026.04.01, Pub. Date: 8 Aug. 2026

This paper presents a hybrid quantum-classical framework for real-time estimation of cognitive engagement from multichannel electroencephalography (EEG) using a new operational indicator called Computational Mental Energy (CME). The proposed approach integrates signal preprocessing (windowing, filtering, spectral feature extraction), spectral feature extraction, a 4-qubit variational quantum classifier (VQC) for flow-state probability estimation, and a metaheuristic optimization loop for balancing predictive quality and quantum resource cost. CME is defined as a window-level function of aggregated spectral energy, task complexity, and estimated flow probability, measured in a dedicated signal-energy unit called Vernik (Vn), with session-level aggregation rules. The system supports quantum-only, classical-only, and hybrid inference modes and is designed for streaming deployment with wearable EEG devices and server-side inference services. A single-subject pilot study involving eight cognitive activities and EEG recordings from a Muse Athena headband demonstrates that the hybrid mode (μ = 0.6) achieves 0.914 AUROC for flow-state detection, compared to 0.548 for the standalone quantum model, while reducing prediction variance by 40.9%. Validation on the IBM Marrakesh 156-qubit Heron r2 quantum processor shows strong agreement between simulator and hardware results (r = 0.869, MAE = 0.045), confirming the practical feasibility of execution on current quantum hardware. Across activities, CME rates differed significantly, with approximately a nine fold gap between coding and resting states, illustrating the framework’s ability to capture activity-dependent cognitive demand. The proposed architecture provides a reproducible pipeline for EEG-based cognitive-state analytics, resource-aware quantum inference, and future adaptive human-computer interaction systems.

[...] Read more.
Development of a Real-Time Maize Leaf Disease Classification System Deployed on Web and Mobile-Based Applications

By Kennedy O. Okokpujie Osondu C. Ronald Joshua S. Mommoh Mary O. Ogundele OluwadamiI Oguntuyo

DOI: https://doi.org/10.5815/ijem.2026.04.02, Pub. Date: 8 Aug. 2026

Agriculture remains at the core of human life, providing staple food and livelihood for millions worldwide. Among its different domains, food crops directly enter the human system, while cash crops are grown primarily for monetary gains. Maize, as one of the most extensively grown and consumed food crops, is of gigantic economic and nutritional value, particularly in West Africa. Unfortunately, maize plant diseases have adversely impacted farmer yields, resulting in decreased maize production. This research aims to create a system that can identify diseases in maize based on images of the leaves. Three deep convolutional neural network (DCNN) models, namely MobileNetV2, InceptionV3, and ResNet50, were selected to achieve this goal because of their prior ability. The transfer learning technique was adopted to develop new models for classifying maize disease using a hybrid maize leaf image dataset comprising 6,543 images from the University of Pretoria and Kaggle repositories. Furthermore, the dataset was split into 80% for training, 10% for validation, and 10% for testing and the three model were configured and trained. According to the evaluation results, MobileNetV2 was the best model for classifying maize leaf diseases, with a 95.29% classification accuracy. In comparison, InceptionV3 and ResNet-50 yielded accuracies of 92.18% and 74.48%, respectively. The MobileNetV2 was chosen for the dual deployment in both a web-based and a mobile application due to its exceptional performance metrics and its lightweight API. Evaluation of the deployed MobileNetV2 model on both the web and mobile applications showed that it achieved an average confidence rate of 91% on both platforms, with response times of 0.159 s and 0.163 s, and throughputs of 5.88 images/s and 6.28 images/s, respectively. This research offers a simple and intuitive tool for users and agricultural professionals to quickly detect maize leaf diseases and take necessary precautions to mitigate losses.

[...] Read more.
Sequential Random Split with Random Feature Subsetting for imbalanced Iron Deficiency Anemia Classification: A Machine Learning Study

By S. N. Lakshmi Malluvalasa Sajana T.

DOI: https://doi.org/10.5815/ijem.2026.04.03, Pub. Date: 8 Aug. 2026

Iron Deficiency Anemia (IDA) is the most common type of anemia and can adversely affect quality of life by reducing oxygen delivery to body tissues. Accurate diagnosis is essential but remains challenging due to the manual interpretation of Complete Blood Count (CBC) parameters, which can be time-consuming and susceptible to human error. Although machine learning models have explored for early IDA detection, imbalanced datasets often lead to biased predictions and reduced performance for minority classes. Furthermore, conventional feature selection approaches may face challenges when dealing with high-dimensional medical data, potentially affecting predictive performance and model robustness. This study proposed a Sequential Random Split (SRS) framework combined with Random Feature Subsetting strategy for iron deficiency anemia (IDA) classification. The framework explores diverse feature combinations and evaluates their impact on predictive performance using a real-world imbalanced IDA dataset. The proposed approach compared with few ensemble learning methods, including Boosting, Stacking, Random Subspace, Random Patches, Extra Trees, Voting, Bagging, and Gradient Boosting. Model performance assessed using Accuracy, Precision, Recall and F1-Score metrics. The proposed SRS framework achieved the highest classification performance among the evaluated ensemble learning approaches. For class-1 (IDA), the model achieved an accuracy of 94.11%, F1-Score of 96%, and Precision of 96%. For Class-0 (non-IDA), the model achieved a Precision of 88%, F1-Score of 88%, and recall of 88%. These findings indicate that the proposed SRS framework provides strong classification performance on the imbalanced IDA dataset. The experimental results indicate that the proposed SRS framework achieved competitive performance for IDA classification on the evaluated dataset. The combination of Sequential random split and Random Feature Subsetting contributed to improved predictive performance compared with the evaluated ensemble learning approaches. These findings suggest that the proposed framework may serve as a useful machine learning strategy for supporting IDA diagnosis, while further validation on larger and more diverse clinical datasets required to assess its generalizability and robustness.

[...] Read more.
Ensemble-Based Modelling for Enhanced Detection of Pneumonia Disease

By Mustafa Oguzhan Ozdemir Kemal Akyol

DOI: https://doi.org/10.5815/ijem.2026.04.04, Pub. Date: 8 Aug. 2026

Pneumonia is a lung condition that is rather prevalent and has the potential to be lethal. The early diagnosis of this disease is absolutely necessary to cut down on the number of fatalities. The aim of this research is to provide a decision-support tool that can help professionals in the field identify cases of pneumonia. The experimental studies used two publicly available datasets in the Kaggle repository. First, experiments were conducted with the pre-trained models. Then, hard and soft voting ensemble learning approaches were implemented using the five most successful deep learning models. According to the results, the soft voting approach outperformed others, with accuracies of 98.55% and 97.26% in two- and three-class datasets, respectively. With this result, a software included this approach has been developed to assist field experts in their decision-making.

[...] Read more.
Hybrid ViT-UNet Framework for Accurate River Segmentation and Buffer Zone Mapping in High-Resolution Satellite Imagery

By T. Satyanarayana Murthy K. Gangadhara Rao Swathi Sowmya Bavirathi

DOI: https://doi.org/10.5815/ijem.2026.04.05, Pub. Date: 8 Aug. 2026

Precise mapping of water bodies is crucial for flood monitoring, disaster risk and response reduction, as well as sustainable water resource management. In this paper, we introduce a deep learning model for effective segmentation of rivers, lakes, and reservoirs from high-resolution Gaofen-2 satellite images. Leveraging the Five-Billion-Pixels dataset-more than 5 billion annotated pixels for 24 land cover classes—our approach solves the problem of segmenting water bodies on various terrains and environmental conditions. The proposed U-Net and ViT-UNet models, with the former employing Vision Transformers to enhance global context perception. For enhancing generalization, the dataset is augmented using Albumentations and flipping, rotation, and scaling transformations. Hybrid loss functions of Dice Loss, Binary Cross-Entropy, and Focal Loss are employed to handle class imbalance, especially for slender river segments. The ViT-UNet model attained 98.8% pixel accuracy, which mirrors its ability to preserve fine detail and large-scale spatial pattern. Mixed-precision training and the AdamW optimizer has enhanced the computational efficiency. Further, demonstrates the potential of transformer-based segmentation models for remote sensing achieved accuracy of 98% for environmental risk management and decision support in disaster-prone areas.

[...] Read more.
Context-Aware Semantic Fusion for Mathematical Formula Equivalence Detection

By Andrii Dyriv Olga Lozynska Victoria Vysotska Dmytro Uhryn Yuriy Ushenko

DOI: https://doi.org/10.5815/ijem.2026.04.06, Pub. Date: 8 Aug. 2026

 The paper investigates the problem of automatically detecting equivalent mathematical formulas in scientific texts. The authors propose a novel hybrid approach that combines structural analysis of formulas (normalisation, Abstract Syntax Tree (AST) construction, and vectorisation) with deep semantic analysis of the surrounding publication text using transformer models such as SciBERT and Sentence Transformers. During the study, a software implementation was developed that uses cosine similarity to assess context proximity and a Siamese neural network for equivalence classification. Evaluated on a custom dataset of 12,500 formula-context pairs from academic papers, the proposed hybrid model achieved an F1-score of 0.88, significantly outperforming baseline models that rely solely on structural (F1: 0.71) or textual (F1: 0.59) features. The experimental results were further visualised using heat maps, dendrograms, and UMAP projections, confirming the model's ability to identify equivalent expressions even when their syntactic notation differs significantly. The proposed framework is promising for use in anti-plagiarism systems, intelligent search services, and digital scientific libraries.

[...] Read more.
Multilingual Fake News Detection Using Machine Learning with Contextual-Based Feature Extraction

By Nikita Garg Pritam Singh Negi

DOI: https://doi.org/10.5815/ijem.2026.04.07, Pub. Date: 8 Aug. 2026

Fake news has become a major challenge in today’s digital environment, particularly in languages where labeled data is limited. Most existing research has primarily focused on English due to the easy availability of annotated datasets, whereas low-resource languages such as Bengali remain underexplored. This study presents a multilingual approach for fake news detection using machine learning with contextual-based feature extraction. The proposed method integrates n-gram techniques with sentence-level contextual embeddings to capture both word-level patterns and semantic meaning. Since labeled data is not available for the Bengali dataset, a translation-based strategy is employed, followed by a pseudo-labeling process to assign labels automatically. The models are trained on English news titles and subsequently evaluated on both English and Bengali datasets to examine their cross-lingual effectiveness. The experimental findings indicate that ensemble-based classifiers such as Random Forest and Gradient Boosting achieve reliable performance across both languages. In some cases, the results for Bengali data are comparable or slightly better than those for English. The study demonstrates that effective fake news detection is possible in low-resource languages using short text data without relying on manually labeled datasets. The proposed approach provides a simple and efficient solution for multilingual fake news detection in data-scarce environments.

[...] Read more.
Enhanced Technique to Find Diabetic Retinopathy

By Punith Kumar M. B. Prashanth kumar A. D. Santhosh Babu K. C. Leela R.

DOI: https://doi.org/10.5815/ijem.2026.04.08, Pub. Date: 8 Aug. 2026

Visual perception relies on the retina, which converts incoming light into interpretable neural information. Diabetic retinopathy (DR), a complication arising from prolonged hyperglycemia, is a major contributor to progressive vision impairment and often remains undetected during its initial stages. The condition manifests in retinal imagery through distinct patterns, including high-intensity and low-intensity lesion regions such as exudates and hemorrhages. This paper proposes an automated framework for simultaneous identification of multiple lesion types in retinal fundus images. The approach begins with image refinement to improve visual quality, followed by intensity-driven segmentation to ex-tract candidate abnormal regions. Descriptive statistical measures—namely mean intensity, variance, standard deviation, and entropy—are computed to characterize these regions and are subsequently utilized as inputs to an Artificial Neural Network (ANN) for classification. To enhance reliability, the method incorporates mechanisms to exclude anatomically similar structures, particularly the optic disc and vascular components, thereby reducing false detections. Evaluation results confirm that the proposed system achieves effective separation between normal and pathological cases, indicating its potential utility in supporting early-stage screening of diabetic retinopathy.

[...] Read more.
A Graph Learning Framework for Analyzing Smart Assistive Sensor Data in Elderly Fall Prediction

By Krishna Kumari R. Padma V.

DOI: https://doi.org/10.5815/ijem.2026.04.09, Pub. Date: 8 Aug. 2026

Falls among older adults represent a critical public health challenge, with approximately 37.3 million fall-related incidents reported globally each year. Early and accurate prediction of falls is essential to enable timely, proactive interventions and to reduce associated injuries and fatalities. This work introduces a graph-based machine learning framework that leverages data from the cStick, a smart assistive Internet of Medical Things (IoMT) device. Bipartite graphs are constructed to model static correlations between multivariate sensor inputs— including heart rate variability (HRV), pressure, distance, SpO2, blood sugar levels, and accelerometer readings— and fall outcomes encoded as no fall, predicted fall, or definite fall. SHAP (SHapley Additive exPlanations) values are further integrated to enhance model interpretability and to identify the most influential sensor features through feature-only graph projections. Kernel Density Estimation (KDE) plots and pairplots are used to visualize feature distributions across fall categories. The proposed framework demonstrates that Pressure and Distance exhibit the strongest correlations with fall decisions (1.000 and −0.946, respectively), providing actionable insights for risk stratification. The integration of graph-based analysis with SHAP interpretability improves both predictive accuracy and transparency, facilitating proactive interventions and enhancing the safety, autonomy, and well-being of elderly individuals in real-world assistive care settings.

[...] Read more.
Optimal Control of the Membrane Module Start-Up Mode in the Membrane Distillation Process

By Lesya Ladieva Roman Dubik Bogdan Korniyenko

DOI: https://doi.org/10.5815/ijem.2026.04.10, Pub. Date: 8 Aug. 2026

 The study considers the issue of optimal control of the start-up mode of the membrane distillation process. The aim of the work is to increase the efficiency of controlling the process of concentrating solutions in a contact membrane distillation unit, which will contribute to reducing the cost and increasing the level of energy saving of the process with prior uncertainty and changes in the permeability of the membrane over time. An analysis of various publications has shown that no single approach has been proposed to control the start-up mode of the membrane module of the process. For control purposes, a mathematical model of the dynamics of the membrane distillation process is proposed. The written mathematical model of the contact membrane distillation process is nonlinear with respect to the temperature of the solution at the outlet of the membrane module, which is also included in the equations that take into account the vapor flow through the membrane. With an increase in the temperature of the solution at the inlet of the membrane module, the temperatures of the solution and distillate at the outlet of the membrane module increase nonlinearly. The optimality criterion was the minimum process start-up time in the presence of restrictions on the final solution temperature. The penalty method and the gradient procedure on the second interval were used to change the control in order to reach the specified regime. The solution depends on the value of the weight coefficients of the penalty functions, which allowed reaching the specified regime in the minimum time.

[...] Read more.
Deep Hybrid Neural Network for Automated Epileptic Seizure Detection from EEG Signals using MAT-LAB

By Swati Chowdhuri Tiyasha Mondal

DOI: https://doi.org/10.5815/ijem.2026.04.11, Pub. Date: 8 Aug. 2026

Epilepsy is a long-term neurological disorder marked by recurring seizures resulting from irregular neuronal activity in the brain. Prompt and precise identification of epileptic events from electroencephalogram (EEG) signals is essential for successful clinical diagnosis. This study introduces a Deep Hybrid Neural Network framework that combines Convolutional Neural Networks (CNN) with the Aquila Optimizer (AO) for the automatic detection of epileptic seizures utilizing EEG data in MATLAB. The suggested system initially converts EEG signals into time–frequency spectrograms through Short-Time Fourier Transform (STFT), allowing the CNN to capture advanced spatial–spectral characteristics. The AO algorithm further improves these features by tuning hyperparameters and choosing the most distinguished fea-ture subsets, thereby boosting classification accuracy and decreasing computational overhead. The Bonn University EEG dataset was used to evaluate the model through a 5-fold cross-validation method, attaining an average accuracy of 95.62%, where per-class sensitivity and specificity surpassed 97%. Comparative evaluation showed that the CNN–AO hybrid sur-passed traditional classifiers in terms of accuracy and convergence reliability. These findings demonstrate the effectiveness of the proposed hybrid framework for automated epileptic seizure detection and suggest its potential suitability for future real-time and wearable healthcare applications following further deployment-oriented validation.

[...] Read more.
An NLP-Based Framework for Fake News Detection Using Contextual and Engineered Features in Communication Technologies

By S.Gopalakrishnan J.Babitha Thangamalar M. Sahaya Sheela M. Mohammed Mustafa Bindu Babu N.Senthil Madasamy

DOI: https://doi.org/10.5815/ijem.2026.04.12, Pub. Date: 8 Aug. 2026

Fake news detection focuses on identifying and preventing the spread of misleading or false information. It is crucial for maintaining the integrity of public discourse and protecting individuals from the harmful effects of misinformation. By ensuring the correctness and reliability of the information, the fake news detection hinders the loss of trust in the media, institutions, and public communication channels. The fake news detection system suggested is in the process of data acquisition where news stories are either manually or automatically retrieved from the net via web crawlers. The collected data later filters the information so it will use only credible sources. Phase two consists of the pre-processing phase using BERT, wherein the data will be tokenized and mapped into contextual embeddings that reflect the semantic meaning of words. Phase three is about engineering features using methods like TF-IDF and Word2vec to 
fine-tune the embeddings and label the important textual features. The final Phase of Classification occurs using the engineered features such that BERT-generated outputs are fine-tuned and passed through softmax functions to ascertain whether the news is fake or real. This holistic and all-encompassing approach integrates advanced natural language processing with feature engineering for an effective system concerning detection of fake news accurately. The model achieved remarkable results over various phases. Training accuracy went from 75% up to those above 95% whereas test accuracy tips above 90%, soaring from below 70%. The model's performance was validated with a balanced confusion matrix and a high ROC AUC of 0.94. Throughout different phases, accuracy, precision, recall, and F1-score increased, reaching 97.0%, 96.7%, 96.8%, and 96.9%, respectively, in the final classification phase, demonstrating robust and reliable detection capabilities.

[...] Read more.
A Comparative Analysis of Conventional Methods for Sensor-Driven Spatial Interpolation for Air Quality Monitoring

By Ajay Kumar P. V. Rajaraman Albins Paul Hasna Hameed

DOI: https://doi.org/10.5815/ijem.2026.04.13, Pub. Date: 8 Aug. 2026

Accurate air quality mapping in regions with sparse sensor deployment remains a challenge due to high infrastructure costs. While modern literature increasingly favors heavy, cloud-based Artificial Intelligence frameworks that assume dense input networks, the operational boundaries and mathematical fidelity of lean, edge-computed spatial interpolation models in ultra-sparse (e.g., 5-node) live frameworks remain poorly defined. This work presents a case-study comparative evaluation of conventional spatial interpolation techniques for estimating pollutant concentrations in sensor-limited environments. Eight interpolation methods, namely Inverse Distance Weighting (IDW), Kriging, Empirical Bayesian Kriging (EBK), Radial Basis Function (RBF), Nearest Neighbor (NN), Akima, Piecewise Cubic Hermite Interpolation (PCHIP), and Cubic Spline (CS), were analyzed using real-time environmental data collected from five locations in Kalady, Kerala, India. The interpolation performance was evaluated using RMSE and MAE metrics through leave-one-out validation. Results indicate that, within this deployment, Kriging and IDW achieved the best estimation accuracy for PM2.5, PM10, humidity, and temperature compared to the other methods evaluated. A LoRa-based sensing framework was also integrated to support low-power real-time environmental monitoring. As a single-region case study based on five monitoring nodes, the findings are specific to this deployment, and broader validation across additional stations and regions is identified as future work. The proposed approach demonstrates the feasibility of resource-efficient air quality estimation in regions with limited monitoring infrastructure.

[...] Read more.
Hybrid Feature Fusion and Bayesian-Optimized Ensemble Learning for Robust Citrus Disease Detection

By Nagineni Venkata Sireesha Gillala Rekha

DOI: https://doi.org/10.5815/ijem.2026.04.14, Pub. Date: 8 Aug. 2026

Detecting citrus diseases at an early stage is very important for ensuring fruit quality and minimizing production losses, as well as for raising awareness about sustainable agriculture. As a solution to this problem, the authors of this paper propose a hybrid feature-based citrus disease classification system that integrates deep learning representations, handcrafted descriptors, feature selection, and ensemble learning, all of which are tuned via Bayesian optimization. We perform tests on two real-world citrus disease datasets that differ greatly in nature: a four-class lemon dataset and a two-class orange dataset. Both datasets were collected under quite different environmental conditions, so they show diverse disease symptoms and varying background complexity. Deep semantic features were obtained by running a pretrained ResNet50 network. In addition to those, other complementary handcrafted features, such as color, texture, spatial, and statistical features, were extracted from the segmented infected areas. Together, the hybrid feature vector of 2082 dimensions was subjected to an optimization method known as Neighborhood Component Analysis (NCA). This technique selects 300 features that are most effective for discrimination while at the same time ensuring the preservation of class separability and the minimization of redundancy. For the classification task, two classifiers, namely Random Forest (RF) and Bayesian-Optimized Random Forest (BORF), were employed. The latter is based on Bayesian optimization to locate the model hyperparameters. To measure the model's performance in an unbiased manner, five-fold cross-validation was performed. Based on the experimental results, BORF can improve classification accuracy on the lemon dataset from 90.42% to 93.75% and on the orange dataset from 95.42% to 95.92% compared to the baseline RF classifier. Cross-validation mean accuracies of the proposed system were 93.75 ± 0.82% and 95.92 ± 0.47% for the lemon and orange datasets, respectively. Receiver Operating Characteristic (ROC) analysis provided class-specific area under the curve (AUC) values of 0.975 and 0.972 for the orange dataset, with a macro-averaged AUC of about 0.94 for the multi-class lemon dataset. The combination of hybrid feature fusion, NCA-based feature optimization, and Bayesian-optimized ensemble classification leads to enhanced discriminative power, greater robustness, and better generalization performance for the system in citrus disease identification in a real-world agricultural setting, as demonstrated by the results.

[...] Read more.
Thermal Emotion Recognition with and without Synthetic Facial Images using Attention-Based EfficientNet5

By Tahir Aman Sintayehu Hirpassa Kibreab Adane

DOI: https://doi.org/10.5815/ijem.2026.04.15, Pub. Date: 8 Aug. 2026

Humans use emotions to express their feelings and effectively interact with others. Humans express emotions through hands, voice, gestures, and, most importantly, facial expressions. Facial emotion recognition is widely used in human-computer interaction, security, and healthcare. Traditional facial emotion recognition using visible light is often affected by changing lighting conditions. Thermal images could be used as an alternative solution because it relies on physiological heat patterns that remain consistent regardless of illumination. However, the use of thermal images for emotion recognition has not been extensively explored, primarily due to the scarcity of thermal image datasets. The study utilized the recently published Thermal Emotion dataset, which contains 2,250 thermal images before augmentation. After generating a synthetic dataset using cGAN, the datasets expanded to 6,823 images. To prevent dataset leakage, the study used 80% of the data for training, 10% for validation, and 10% for testing. The study used a bilateral filter to reduce noise while keeping relevant edge information, used CLAHE to enhance local contrast in low-intensity regions, making subtle thermal gradients more distinguishable, used Gaussian smoothing to reduce high-frequency noise, resulting in more stable feature extraction, and used attention mechanisms, CBAM, to allow the model to focus on emotion-relevant facial regions through its channel and spatial attention components.  The combination of these techniques improved inter-class separability and contributed to measurable gains in recognition accuracy. The study findings show that inclusions of the bilateral filter, CLAHE cGAN, EfficientNetB5, and the CBAM attention mechanism significantly improved accuracy from 97.04% to 98.81%.  For each experiment, ResNet18 classifies human facial emotions into five expressions: Happy, Sad, Angry, Natural, and Surprise. These findings imply that synthetic data generation using a cGAN to overcome data scarcity and an attention mechanism provides robust, lightning-independent solutions for facial emotion recognition.

[...] Read more.
An Interpretable Machine Learning Framework for Breast Cancer Diagnosis Using Statistical Feature Analysis and Ensemble Classification

By T. Haripriya M.V. Ramana Murthy Ch. Vasavi Swathi Gowroju G. Srinivas Devineni Gireesh Kumar

DOI: https://doi.org/10.5815/ijem.2026.04.16, Pub. Date: 8 Aug. 2026

A clear, statistically sound, yet easily understandable breast cancer diagnosis is a difficult issue in all healthcare systems, because early stages of breast cancer are critical in therapy success and long-term survivability. This machine-learning-based breast cancer classifier, in a statistically justified, rigorously experimentally validated way, classifies a set of 569 breast cancer cases with 9 cytological features for breast cancer diagnosis. The classifier uses a rigorous set of data cleanup measures, including missing-value substitution, correlation-based feature reduction, and projection into principal component space, to achieve high data quality, reduce redundancy, and enhance feature usefulness. Five supervised classifiers, in a widely accepted train-test model using an 80:20 random sample split and 5-fold cross-validation, are fitted and evaluated using Accuracy, Precision, Recall, F1-score, and Area under the ROC curve. In these tests, the Random Forest classifier got the best result, with 95.84% Accuracy, 95.31% Precision, 95.12% Recall, 95.21% F1-score and 0.982 area under the ROC curve; in a statistically sound consistency test using cross-validation, its mean accuracy reached 95.96% with a small standard deviation of 0.43. To provide a clear, interpretable indication of which features truly matter, we performed a feature-importance analysis on the best classifier, the Random Forest model. Results show that the expression levels of Bland Chromatin, Single Epithelial Cell Size, Normal Nucleoli, Uniformity of Cell Shape, Uniformity of Cell Size and Bare Nuclei are closely related to breast cancer diagnosis; this is almost the same as the clinical diagnosis findings, and very naturally suggests that abnormalities of cellular morphology and nuclei are major symptoms of breast cancer. In comparison, prior research may neglect validation and efficiency comparisons or focus only on the classifier's accuracy. Our method combines multiple levels of assessment (statistical data-by-data validation, feature importance, cross-validation, and comparison of different classifiers using ensemble learning) into a single evaluation system. This combined approach not only enhances predictive capability but also makes the entire setup more explicitly interpretable from a clinical perspective, thereby making it more suitable for health care decision support. Given the strong classification performance, interpretability, and validation suggested above, the model would help physicians detect breast cancer very early, reducing the risk of misdiagnosis.

[...] Read more.
Design and Implementation of a 16-Bit ALU Using Structural Verilog

By Punith Kumar M. B. Yashaswini H. A.

DOI: https://doi.org/10.5815/ijem.2026.04.17, Pub. Date: 8 Aug. 2026

The Arithmetic Logic Unit (ALU) is a fundamental building block of all modern central processing units (CPUs). This papers  presents the design and implementation of a 16-bit, 4-function ALU constructed entirely from basic logic gates, using a structural Verilog approach. The ALU supports four operations: addition, subtraction, bitwise AND, and bitwise OR, implemented via a modular, bit-sliced architecture. A single ripple-carry adder with 2's complement logic performs both arithmetic operations efficiently. Operation selection is achieved through a 2-bit control signal and a 4-to-1 multiplexer per bit slice. The system was verified using QuestaSim simulation with both normal and boundary test vectors. Unlike existing behavioral or high-level implementations, this work contributes a fully gate-level structural model that makes every interconnection explicit, enabling transparent analysis of carry propagation and delay. Performance metrics including propagation delay and hardware resource usage are analyzed. The results validate the design's functional correctness and demonstrate its potential as a scalable building block for processor architectures.

[...] Read more.
Modelling and Intellectual Analysis of Quantitative Characteristics of Air Pollution

By Volodymyr Hura

DOI: https://doi.org/10.5815/ijem.2026.04.18, Pub. Date: 8 Aug. 2026

The current state of environmental safety requires the introduction of the latest technologies for monitoring and analyzing environmental data. Air pollution with fine particles (PM2.5, PM10) from local emission sources creates significant computational challenges due to insufficient data and the dynamic nature of pollution propagation processes.
Objective. The goal of the work is to solve these problems by developing and integrating modern methods and tools for modeling and intelligent analysis of air pollution characteristics. A prototype hybrid algorithmic pipeline is proposed that integrates an adapted Gaussian model with optimized machine learning and neural network models. The method utilizes a Mamdani-type fuzzy logic system to determine the Atmospheric Stability Class based on continuous meteorological inputs. Additionally, Bayesian inverse modeling using Markov Chain Monte Carlo (MCMC) methods is applied to estimate unknown source emission intensity. The approach is implemented using cloud technologies (Azure Data Lake) and edge computing systems (Nvidia Jetson Nano). A large-scale comparative analysis of deep neural network architectures (Bidirectional LSTM, CNN) and ensemble models (XGBoost, CatBoost) was conducted. The Bidirectional LSTM provided the best overall performance (MSE=0.521, R2=0.985). The integration of fuzzy stability inputs reduced the MSE by 16%. The experiments conducted and numerical modeling confirmed the effectiveness of the proposed methods and the operability of the developed neuro-controller system. The results allow recommending the integrated approach for real-time environmental monitoring and decision support in data-scarce environments.

[...] Read more.
Deep Learning Models for Non-Invasive Blood Group Detection Using Fingerprint Images

By Rajendra Prasad Banavathu James Stephen Meka K. Venkateswara Rao B. Raja Rao Yaswanth Kumar Peddagamalla

DOI: https://doi.org/10.5815/ijem.2026.04.19, Pub. Date: 8 Aug. 2026

In the context of medical diagnosis, the identification of human blood groups plays a significant role. To perform the identification of human blood groups, usually invasive identification methods are used. However, due to the limitations of the invasive methods of blood group identification, the use of fingerprint-based identification of human blood groups gained significance recently. In this paper, an accurate comparison of the recently developed deep learning models of fingerprint-based blood group identification techniques is provided. Five CNN-based models, such as ConvNeXt-Tiny, ConvNeXt-Small, EfficientNetV2-S, RepVGG-B3g4, and MobileViT-V2 models for the identification of human blood groups, are implemented. The results obtained in the experiment, considering the accuracy of the models, have proved the EfficientNetV2-S model to have the highest accuracy of 98.84%, followed by the ConvNeXt-Small, ConvNeXt-Tiny, MobileViT-V2, and RepVGG-B3g4 models with accuracies of 97.66%, 97.21%, 97.08%, and 96.98%, respectively.

[...] Read more.
Software for Simulating Neural Network Workload Impact on Mobile Devices

By Kirill Smelyakov Oleksandr Dolhanenko Oleksiy Lanovyy Victoria Vysotska Dmytro Uhryn

DOI: https://doi.org/10.5815/ijem.2026.04.20, Pub. Date: 8 Aug. 2026

Mobile processors are now capable of running complex machine learning tasks locally, yet our mobile operating systems often hold them back. To preserve battery life, heavy workloads are typically restricted to charging or idle periods, severely limiting time-sensitive applications such as real-time health monitoring or federated learning. While we need more innovative scheduling algorithms to overcome this, validating them on physical hardware is extremely difficult. Factors like thermal throttling, background kernel activity, and battery degradation create an unpredictable environment where no two tests are ever quite the same. To solve this reproducibility challenge, we introduce a new simulation framework. Instead of relying on inconsistent test runs on physical devices, our system records a device's natural baseline activity and mathematically superimposes the resource footprint of a heavy task. It allows us to model CPU saturation, energy loss, memory pressure, and thermal dynamics in a controlled software environment. When compared against ground-truth recordings from Samsung Galaxy S10e and Fold 5 devices, the simulator achieved correlation scores exceeding 0.90 for computing and thermal metrics. By isolating the workload's impact from environmental noise, this platform provides a scalable method for benchmarking neural network schedulers without the logistical bottlenecks associated with continuous hardware testing.

[...] Read more.
A Novel Explainable LLM-based Why-QA Framework for Climate Resilient and Sustainable Smart Agriculture

By Manvi Breja

DOI: https://doi.org/10.5815/ijem.2026.04.21, Pub. Date: 8 Aug. 2026

Sustainable smart agriculture and climate resilience are significant factors for maintaining food security with environmental management. Existing agricultural systems target predicting and generating reports but lack detailed explanations as to why the phenomenon occurs. The examples of such why are like “Why there is a significant decline in the yield?”, “Why the soil is degrading?”, “Why the level of water is declining?” and so on. To address this challenge, the paper presents a prototype for Explainable LLM based Why-QA framework for sustainable smart agriculture. The implemented prototype integrates domain-enriched LLM with knowledge graph, causal inference engine with explainability layer to provide detailed explanations to complex “why-questions”. Domain-specific LLMs are used to support the domain knowledge with knowledge graph analyzing the sustainable relationships in the answer, causal reasoning to produce the causes of events and explainability module to provide the detailed reasoning supported with benchmark sustainable metrics incorporating the soil, climate and crop data. The prototype implementation of framework is evaluated on a dataset of 500 annotated agricultural why-type questions constructed from FAOSTAT, USDA, and NOAA sources. Results clearly demonstrate promising improvements over five baseline systems developed across causal reasoning and explanation quality metrics, which help validating the architectural feasibility of the proposed framework. 

[...] Read more.
FL-EZTF: A Privacy-Preserving Federated Deep Learning Framework with Enhanced Zero Trust for Healthcare IoT Security

By Urvashi Parul Agarwal Kamlesh Kumar Raghuvanshi Jawed Ahmed

DOI: https://doi.org/10.5815/ijem.2026.04.22, Pub. Date: 8 Aug. 2026

Smart healthcare IoT systems are vulnerable to cyber threats as they deal with sensitive patient information. Problems such as privacy, scalability, and delayed response to threats in distributed healthcare environments challenge centralized security approaches. To mitigate the security challenges of cloud-edge healthcare IoT systems, this paper presents FL-EZTF, a privacy-preserving, Federated Deep Learning and Enhanced Zero Trust Framework. The framework combines federated learning, Enhanced Zero Trust Architecture (E-ZTA), and Secure Access Service Edge (SASE). In this framework, lightweight deep learning models are developed locally at hospitals and various edge nodes without the need to transfer sensitive medical data. In place of raw data, model updates are sent conveniently through a trustaware federated learning process. Simultaneously, E-ZTA performs continuous authentication, micro-segmentation, and access control to rapidly contain threats. The framework is assessed using CIC-IoT-2023, IoT-23, and WESAD datasets. The experimental results show improved accuracy in detection, lower rates of false positives, a significant reduction in the latency of decisions, and enhanced containment as compared to centralized and traditional federated learning.

[...] Read more.
AuthProtect: An Incremental Learning Framework for Android Malware Detection via Permission-Exploitation Mapping

By Maksim Iavich Razvan Bocu

DOI: https://doi.org/10.5815/ijem.2026.04.23, Pub. Date: 8 Aug. 2026

Android's widespread adoption and open ecosystem make it a primary target for malware, a challenge exacerbated by internet fragmentation resulting in non-stationary data distributions across regions. This work presents AuthProtect, a scalable malware detection framework based on incremental learning and a novel permission-to-exploitation mapping approach that links 135 permissions to 25 malware development techniques. The system is validated on a balanced dataset of 82,704 benign and 82,704 malicious applications, partitioned into three geographic regions to assess robustness to distribution shifts. A similarity-based selective training strategy improves computational efficiency by training only on novel samples (cosine similarity < threshold τ), while a test-then-train mechanism enhances robustness by sequentially processing samples to avoid data exposure bias. Evaluation on four benchmark datasets (Naticusdroid, Malgenome, CICMalDroid 2020, Android Malware Dataset) demonstrates accuracy ranging from 0.9573 to 0.9992, with a maximum accuracy of 0.9982 on real-world data. We provide a comparative analysis against state-of-the-art methods and ablation studies quantifying the contribution of each component. Limitations include dependency on the completeness of permission-technique mapping and computational overhead for real-time deployment on resource-constrained devices.

[...] Read more.
Performance Evaluation of Bagging and Boosting- Based Ensemble Learning Models for Clinical Liver Disease Prediction Using the LDPD Dataset

By A. S. M. Shafi

DOI: https://doi.org/10.5815/ijem.2026.04.24, Pub. Date: 8 Aug. 2026

The liver is one of the most essential internal organs in the human body, acting as a metabolic powerhouse and playing a key role in the immune system. However, Liver Diseases (LD) are rising globally, driven by unhealthy lifestyles and excessive alcohol use. Liver diseases cause millions of deaths annually all over the world, highlighting the need for early diagnosis. This study aims to evaluate the implication of ensemble machine learning techniques—bagging and boosting—for liver disease prediction, utilizing 30,691 instances and 11 features of the Liver Disease Patient Dataset (LDPD). To improve model performance, hyperparameter tuning, outlier removal, normalization for data scaling and feature importance to identify the most significant predictors are used. Eight state-of-the-art ensemble models were evaluated in this study: Random Forest (RF), Extra Trees Classifier (ETC), Bagged Decision Tree (Bagged DT), Adaptive Boosting (AdaBoost), Gradient Boosting (GradientBoost), Xtreme Gradient Boosting (XGBoost), Categorial Boosting (CatBoost) and Light Gradient Boosting Machine (LightGBM). Our experimental results showed that RF algorithm outperformed other algorithms, achieving the highest accuracy, specificity, precision, and F1-score of 99.85%, 99.85%, 99.93% and 99.89% respectively. While LightGBM attained the highest recall rate of 99.86% making it particularly suitable for identifying true positive cases and minimizing the missed diagnosis. These findings highlight the effectiveness of ensemble learning methods (bagging and boosting algorithms) in accurately predicting liver disease.

[...] Read more.
Deep Convolutional Auto encoders with Structured Latent Space for Fashion-MNIST Denoising and Classification

By Pattapu Sravani P. Satya srinivasa babu Inturu Bhavani Siva Phanindra Shaik Hasane Ahammad Ramachandran Thandaiah Prabu Ahmed Nabih Zaki Rashed

DOI: https://doi.org/10.5815/ijem.2026.04.25, Pub. Date: 8 Aug. 2026

Fashion image analysis has a lot of trouble working with noisy data, especially when there aren't any good training examples to use. This happens because standard auto encoders can't keep class separability when the noise level changes. This paper introduces Structured Auto encoders (SAE), which combine convolutional encoder-decoder architectures with geometric constraints in the latent space. This strategy uses greedy layer-wise pre-training, and then it fine-tunes the model using two loss functions: structured latent space regularization and reconstruction error. The encoder has convolutional layers with 3×3 filters and ReLU functions that only turn on when they are needed. During testing, it was tested on Fashion-MNIST with noise levels ranging from 20% to 70%. It was also tested on MNIST, DeepFashion2, and 3D human pose datasets. The results of the experiment show that SAE can classify 85.4% of the time with only five samples labeled for each group. This is much better than the standard auto encoders (68.4%) and the baseline support vector machine (SVM) methods (84.32%). With the best setup, which has 1,000 hidden nodes and a learning rate of 0.1, images with up to 70% noise can be reconstructed well With 80% classification accuracy on clean test data, the latent space structured method allows for the solution of basic issues that are related to introducing geometric constraints between the classes so that the class boundaries remain intact even in a noisy environment.The proposed method achieves a classification accuracy of 85.4% ± 1.2%under standard evaluation settings, with robustness validated under noise levels of up to 70%. All results are obtained using standard training–testing splits and are averaged over multiple experimental runs to ensure reliability and reproducibility. The proposed semi-supervised performance is enhanced by increasing the number of discriminative features through the use of a highly structured representation of the data, as opposed to using an unstructured representation, which results in performance enhancement. The author presents an original framework capable of simultaneously providing robust image classification and denoising capability through geometric constraints established in the framework, which enhances the learning of discriminative features through the use of limited amounts of labeled data. The implementation of this framework in the real-world fashion industry would facilitate the processing of fashion images with respect to noise resistance and the ability to annotate images quickly.

[...] Read more.