IJEM Vol. 16, No. 5, 8 Oct. 2026
Cover page and Table of Contents: PDF (size: 534KB)
PDF (534KB), PP.227-237
Views: 0 Downloads: 0
Programmable Logic, System on a Chip, Matrix Multiplication, Convolutional Neural Networks, Denoising, Deep Learning
The contemporary challenges facing computer engineering have led to a focus on improving the efficiency of hardware and computing systems, as well as on optimising the operation of devices for effective data processing. In most modern systems of artificial intelligence, computer vision and digital signal processing, matrix multiplication is a basic operation. With the increasing resolution of sensors and the growing complexity of neural networks, classical general-purpose processors face the ‘von Neumann bottleneck’, where the data transfer rate between memory and the processor is limited, and the sequential execution of instructions does not allow the required real-time throughput to be achieved. The subject of this research is a System-on-a-Chip architecture that combines a dual-core processor based on the ARM architecture with programmable logic. This hybrid structure allows the most computationally intensive tasks to be offloaded to the hardware, whilst leaving control and the implementation of high-level interfaces to software. A neural network accelerator based on programmable logic devices offers advantages such as the capability for stream processing, which minimises the number of accesses to external memory, and the ability to utilise so-called mixed-precision computing. Furthermore, the accelerator’s efficiency is achieved through the use of multiple processing elements, enabling parallel computation of the neural network’s output channels. It has been demonstrated that implementing matrix operations on programmable logic enables parallelism of hundreds of operations per clock cycle, which is unattainable for systems based on general-purpose CPUs of the same class. The proposed implementations for a programmable logic-based accelerator have demonstrated speed-up compared to a system based on an ARM-architecture processor, as well as high adaptability to Deep Learning tasks.
Yaroslav Klyatchenko, Oxana Tarasenko Klyatchenko, "The Efficiency of FPGA-based Matrix Multiplication Accelerators for Neural Network Algorithms", International Journal of Engineering and Manufacturing(IJEM), Vol.16, No.5, pp. 227-237, 2026. DOI:10.5815/ijem.2026.05.12
[1]Girish D. Chate, S. S. Bhamare, "Classification of Soil Images Using Convolutional Neural Network", International Journal of Image, Graphics and Signal Processing, Vol.17, No.5, pp. 26-41, 2025. DOI:10.5815/ijigsp.2025.05.03.
[2]L. Yasenko, Y. Klyatchenko and O. Tarasenko-Klyatchenko, "Image noise reduction by denoising autoencoder," 2020 IEEE 11th International Conference on Dependable Systems, Services and Technologies (DESSERT), Kyiv, Ukraine, 2020, pp. 351-355, doi: 10.1109/DESSERT50317.2020.9125027.
[3]Comparison Analysis of Cycles X Render and Cycles Render Using Google Colab. (2023). Jurnal Tika, 8(1), 90-94. https://doi.org/10.51179/tika.v8i1.1937.
[4]Yasenko, L., & Klyatchenko, Y. (2021). Convolutional properties of a neural network based on autoencoders. Information Technologies and Computer Engineering, 18(3), 77-85. https://doi.org/10.31649/1999-9941-2021-52-3-77-85.
[5]Ghimire, D., Kil, D., & Kim, S.-h. (2022). A Survey on Efficient Convolutional Neural Networks and Hardware Acceleration. Electronics, 11(6), 945. https://doi.org/10.3390/electronics11060945.
[6]Gonzalez R.C., Woods R.E. Digital image processing. 2nd ed. Upper Saddle River, N.J: Prentice Hall, 2002. 793 p.
[7]Maladkar, K.: 6 Types of Artificial Neural Networks Currently Being Used in Machine Learning. Analytic Indian Magazine, 15 Jan 2018. https://analyticsindiamag.com/6-types-of-artificial-neural-networks-currently-being-used-in-todays-technology/. Accessed 9 Oct 2020.
[8]P. Bharati and A. Pramanik, "Deep learning techniques—R-CNN to mask R-CNN: a survey," in Proc. Int. Conf. Comput. Intell. Pattern Recognit. (CIPR), 2019, pp. 657–668.
[9]K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
[10]Kondru Mounika, P. Venkateswara Rao, Anand Anbalagan, "Modified CNN Model for Network Intrusion Detection and Classification System Using Local Outlier Factor-based Recursive Feature Elimination", International Journal of Computer Network and Information Security, Vol.17, No.1, pp.82-91, 2025. DOI:10.5815/ijcnis.2025.01.07
[11]Devesh Kumar Srivastava, Chirag Goel, K. Kishore Kumar, Akhilesh Kumar Sharma, Babu R. Dawadi, Eshaan Saha, "Enhancing Underwater Object Detection through CNN-based Image Enhancement and Classification", International Journal of Engineering and Manufacturing (IJEM), Vol.16, No.2, pp.91-110, 2026. DOI:10.5815/ijem.2026.02.06
[12]L. Chen, S. Li, Q. Bai, J. Yang, S. Jiang, and Y. Miao, “Review of image classification algorithms based on convolutional neural networks,” Remote Sensing, vol. 13, no. 22, p. 4712, 2021.
[13]Yilin Zhao Systematic Analysis on TPU’s History and Key Technologies Behind It // Proceedings of the 2025 International Conference on Electronics, Electrical and Grid Technology (ICEEGT2025), Advances in Engineering Research 292, https://doi.org/10.2991/978-94-6463-986-5_46/
[14]Takano S. Thinking Machines. Machine Learning and Its Hardware Implementation. Academic Press, 2021. 306 p.
[15]N. Skilkov and Y. Klyatchenko, "Complex methodology for determining WCET in modeling real-time system operation processes," 2024 14th International Conference on Dependable Systems, Services and Technologies (DESSERT), Athens, Greece, 2024, pp. 1-6, doi: 10.1109/DESSERT65323.2024.11122223.
[16]Chellappa, R., Theodoridis, S. Academic Press Library in Signal Processing, Volume 7: Array, Radar and Communications Engineering. — Elsevier, 2021.
[17]Matrix Multiplication Background User Guide: NVIDIA Deep Learning Performance Documentation. https://docs.nvidia.com/deeplearning/performance/dl-performance-matrix-multiplication/index.html.
[18]Klyatchenko Ya.M., Yasenko L.S. COMPARATIVE ANALYSIS OF THE EXECUTION SPEED OF NEURAL NETWORK ALGORITHMS DEPENDING ON HARDWARE. // APPLIED MATHEMATICS AND COMPUTING XIV Conference of Young Scientists PMK-2021. Kyiv: "PROSVITA", 2021. https://ela.kpi.ua/handle/123456789/55884.
[19]Zynq 7000 SoC Technical Reference Manual (UG585). https://docs.amd.com/r/en-US/ug585-zynq-7000-SoC-TRM.
[20]Yang Zhang Arm Neon programming quick reference. Arm NEON programming quick reference guide. https://developer.arm.com/community/arm-community-blogs/b/operating-systems-blog/posts/arm-neon-programming-quick-reference.
[21]Zynq-7000 SoC Product Selection Guide (XMP097) URL: https://docs.amd.com/v/u/en-US/zynq-7000-product-selection-guide.
[22]A Zynq Accelerator for Floating Point Matrix Multiplication Designed with Vivado HLS (XAPP1170). https://docs.amd.com/v/u/en-US/xapp1170-zynq-hls.
[23]AMD Adaptive Computing. 7 Series FPGAs DSP48E1 Slice User Guide (UG479). San Jose, CA: AMD, 2022.
[24]Xilinx Inc. Zynq-7000 SoC Technical Reference Manual (UG585). San Jose, CA: Xilinx, 2023. 1850 p.
[25]AMD Vivado™ High-Level Design URL: https://www.amd.com/en/products/software/adaptive-socs-and-fpgas/vivado/high-level-design.html#resources.
[26]Xilinx Inc. Vitis High-Level Synthesis User Guide (UG1399). San Jose, CA: AMD, 2024. 320 p.
[27]C. Jiang, D. Ojika, B. Patel, and H. Lam, “Optimized fpga-based deep learning accelerator for sparse cnn using high bandwidth memory,” in Proceedings of FCCM, 2021, pp. 157–164.
[28]Zhang, C., Li, P., et al. Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks. — Proceedings of the 2015 ACM/SIGDA International Symposium on FPGAs, pp. 161–170.