Context-Aware Semantic Fusion for Mathematical Formula Equivalence Detection

PDF (940KB), PP.81-96

Views: 0 Downloads: 0

Author(s)

Andrii Dyriv 1,* Olga Lozynska 1 Victoria Vysotska 1 Dmytro Uhryn 2 Yuriy Ushenko 2

1. Information Systems and Networks Department, Lviv Polytechnic National University, Lviv, 79013, Ukraine

2. Department of Computer Science, Yuriy Fedkovych Chernivtsi National University, Chernivtsi, 58012, Ukraine

* Corresponding author.

DOI: https://doi.org/10.5815/ijem.2026.04.06

Received: 28 Aug. 2025 / Revised: 25 Feb. 2026 / Accepted: 28 Mar. 2026 / Published: 8 Aug. 2026

Index Terms

Mathematical formulas, semantic analysis, equivalence of formulas, publication context, vector representation (Embeddings), transformer models (SciBERT, MathBERT), machine learning, clustering, natural language processing (NLP)

Abstract

 The paper investigates the problem of automatically detecting equivalent mathematical formulas in scientific texts. The authors propose a novel hybrid approach that combines structural analysis of formulas (normalisation, Abstract Syntax Tree (AST) construction, and vectorisation) with deep semantic analysis of the surrounding publication text using transformer models such as SciBERT and Sentence Transformers. During the study, a software implementation was developed that uses cosine similarity to assess context proximity and a Siamese neural network for equivalence classification. Evaluated on a custom dataset of 12,500 formula-context pairs from academic papers, the proposed hybrid model achieved an F1-score of 0.88, significantly outperforming baseline models that rely solely on structural (F1: 0.71) or textual (F1: 0.59) features. The experimental results were further visualised using heat maps, dendrograms, and UMAP projections, confirming the model's ability to identify equivalent expressions even when their syntactic notation differs significantly. The proposed framework is promising for use in anti-plagiarism systems, intelligent search services, and digital scientific libraries.

Cite This Paper

Andrii Dyriv, Olga Lozynska, Victoria Vysotska, Dmytro Uhryn, Yuriy Ushenko, "Context-Aware Semantic Fusion for Mathematical Formula Equivalence Detection", International Journal of Engineering and Manufacturing (IJEM), Vol.16, No.4, pp.81-96, 2026. DOI:10.5815/ijem.2026.04.06

Reference

[1]Sakshi, Kukreja, V. "Image Segmentation Techniques: Statistical, Comprehensive, Semi-Automated Analysis and an Application Perspective Analysis of Mathematical Expressions," Arch. Comput. Methods Eng. 30(1), 457–495 (2023). https://doi.org/10.1007/s11831-022-09805-9.
[2]Schmitt-Koopmann, F.M., Huang, E.M., Hutter, H.P., Stadelmann, T., Darvishy, A. "FormulaNet: A Benchmark Dataset for Mathematical Formula Detection," IEEE Access 10, 91588–91596 (2022). https://doi.org/10.1109/ACCESS.2022.3202639.
[3]Li, Z., Wu, Y., Li, Z., Wei, X., Yang, F., Zhang, X., Ma, X. "Autoformalize Mathematical Statements by Symbolic Equivalence and Semantic Consistency," In: Advances in Neural Information Processing Systems (NeurIPS) 37, pp. 53598–53625 (2024). https://doi.org/10.52202/079017-1697.
[4]Niu, K., Zhang, P. The Mathematical Theory of Semantic Communication. Springer, Singapore, pp. 1–52 (2025). https://doi.org/10.1007/978-981-96-5132-0.
[5]Peng, S., Yuan, K., Gao, L., Tang, Z. "MathBERT: A Pre-Trained Model for Mathematical Formula Understanding," arXiv 2105.00377 (2021). https://doi.org/10.48550/arXiv.2105.00377.
[6]Ma, W., Liu, S., Lin, Z., Wang, W., Hu, Q., Liu, Y., Zhang, C., Nie, L., Li, L., Liu, Y. "LMs: Understanding Code Syntax and Semantics for Code Analysis," arXiv 2305.12138 (2023). https://doi.org/10.48550/arXiv.2305.12138.
[7]Wagh, V., Laddha, S., Kadam, P. "Detecting Plagiarism Using Latent Semantic Analysis and Cosine Similarity Approach," In: 2024 IEEE International Conference on Blockchain and Distributed Systems Security (ICBDS), pp. 1–6 (2024). https://doi.org/10.1109/ICBDS61829.2024.10837475.
[8]Satpute, A., Greiner-Petter, A., Gießing, N., Beckenbach, I., Schubotz, M., Teschke, O., et al. "Taxonomy of Mathematical Plagiarism," In: European Conference on Information Retrieval (ECIR), pp. 12–20 (2024). https://doi.org/10.1007/978-3-031-56066-8_2.
[9]Ali, A., Taqa, A.Y. "Analytical Study of Traditional and Intelligent Textual Plagiarism Detection Approaches," J. Educ. Sci. 31(1), 8–25 (2022). https://doi.org/10.33899/edusj.2021.131895.1192.
[10]Scharpf, P., Schubotz, M., Cohl, H.S., Breitinger, C., Gipp, B. "Discovery and Recognition of Formula Concepts Using Machine Learning," Scientometrics 128(9), 4971–5025 (2023). https://doi.org/10.1007/s11192-023-04667-9.
[11]Vysotska, V. "Linguistic Intellectual Analysis Methods for Ukrainian Textual Content Processing," In: CEUR Workshop Proceedings 3722, 490–552 (2024). Available: https://ceur-ws.org/Vol-3722/paper25.pdf.
[12]Fedchuk, R., Vysotska, V. "Mathematical Model of a Decision Support System for Identification and Correction of Errors in Ukrainian Texts Based on Machine Learning," In: CEUR Workshop Proceedings 4005, 29–50 (2025). Available: https://ceur-ws.org/Vol-4005/paper3.pdf.
[13]Kholodna, N., Vysotska, V., Markiv, O., Chyrun, S. "Machine Learning Model for Paraphrases Detection Based on Text Content Pair Binary Classification," In: Proceedings of MoMLeT+DS, pp. 283–306 (2022). Available: https://ceur-ws.org/Vol-3312/paper23.pdf.
[14]Mansouri, B., Rohatgi, S., Oard, D.W., Wu, J., Giles, C.L., Zanibbi, R. "Tangent-CFT: An Embedding Model for Mathematical Formulas," In: Proceedings of the 2019 ACM SIGIR International Conference on Theory of Information Retrieval (ICTIR), pp. 11–18 (2019). https://doi.org/10.1145/3341981.3344235.
[15]Nahar, K.M., Alshtaiwi, M.M., Alikhashashneh, E., Shatnawi, N., Al-Shannaq, M.A.A., Abual-Rub, M., BaniIsmail, B. "Plagiarism Detection System by Semantic and Syntactic Analysis Based on Latent Dirichlet Allocation Algorithm," Int. J. Adv. Soft Comput. Its Appl. 16(1) (2024). https://doi.org/10.15849/IJASCA.240330.03.