Nataliia Tytova

Work place: Dragomanov Ukrainian State University, Kyiv, 01601, Ukraine

E-mail: n.m.tytova@udu.edu.ua

Website:

Research Interests:

Biography

Nataliia Tytova: Doctor of Pedagogical Sciences (2019), Professor (2021), Head of Department of Department of Professional Training, Document Studies and Public Administration Dragomanov Ukrainian State University.
Research Interests: psychological and pedagogical training of vocational education teachers, professional training of specialists in the field of document studies, public management and administration.

Author Articles
Specialized Software Tools for Educational Data Mining of Technical Students: Image Segmentation, Clustering, and Correlation-Regression Analysis

By Oleksandr Derevyanchuk Serhiy Balovsyak Shaocheng Qu Yurii Ushenko Nataliia Tytova Hanna Kravchenko

DOI: https://doi.org/10.5815/ijmecs.2026.05.05, Pub. Date: 8 Oct. 2026

This article presents a set of specialized software tools for educational data mining and describes the integration of the developed tools into the educational process of students of technical specialties. The developed set consists of three programs that implement the main tasks of educational data mining. Program No. 1 provides segmentation, analysis, and visualization of images of educational materials; Program No. 2 performs data clustering; Program No. 3 performs correlation and regression analysis to identify relationships and predict learning outcomes. The developed programs can be used to implement a systematic analysis of education quality, during which automated processing of educational data and the formation of recommendations for improving the educational process are performed. The programs are implemented in Python. A scheme for integrating specialized programs into the educational process has been developed. The scheme reflects the relationships between the programs that process educational materials, analyze the results of the educational process, influence students’ educational trajectories and the educational process.
Program No. 1, “SegmentFuzzy24,” is designed for segmentation, analysis, and visualization of images of educational materials. Such processing makes it possible to highlight the investigated elements of technical devices in order to focus students’ attention, as well as to determine the quantitative and geometric characteristics of multimedia presentation slides, including the number and size of symbols, the area of graphical objects, etc. Segmentation is performed using the region growing method. The determined characteristics are used for the classification of educational materials or are transferred to Programs No. 2 and No. 3 for further cluster, correlation, and regression analysis.
Program No. 2, “ClusterFuzzy23,” is designed for clustering educational data, including students’ learning outcomes and the characteristics of educational materials obtained from Program No. 1. Objects are divided into groups with similar characteristics using the K-means method. During the clustering process, the sum of the squared Euclidean distances between each object and the centroid of its cluster is minimized. Data clustering ensures their differentiated processing and contributes to the adjustment of educational trajectories. The clustering results can be transferred to Program No. 3 for further analysis. Fuzzy membership functions are used to determine the degree of membership of objects in overlapping clusters.
Program No. 3, “CorrelRegres25,” is designed for correlation and regression analysis of educational data obtained directly or as a result of clustering in Program No. 2. Correlation analysis is used to establish relationships between the characteristics of the educational process, while regression analysis is applied to model their dependencies using polynomial functions and to predict learning outcomes. The obtained models are used to adjust educational trajectories, while objects that have significant deviations from regression dependencies are classified as outliers.
The developed programs have been demonstrated on sample data to process real educational data, namely for image segmentation and visualization, clustering of educational data, and their correlation and regression analysis.
Data clustering and correlation–regression analysis were performed using the grades of N0 = 76 students in 12 subjects, assessed on a 100-point scale, for courses completed during the first and second years of study. The initial sample of N0 students was randomly divided into a training set (N = 60, 80% of the data) and a validation set (NV = 16, 20% of the data). The quality of the clustering results was evaluated using the Silhouette coefficient, Calinski–Harabasz index and Davies–Bouldin index. The analysis indicated that three or four clusters provide an appropriate representation of the data. The optimal polynomial degree pA = 3 for the regression model was determined based on the minimum root mean square error RmseV = 12.37458, obtained on the validation set. The relatively high RmseV value can be attributed to the influence of a considerable number of factors affecting students’ academic performance. The resulting regression model was used to predict students’ grades in a subject for the subsequent academic period based on their grades in the same subject during the previous period. A promising direction for further improvement of the developed prediction system is the integration of artificial neural networks.

[...] Read more.
Other Articles