Oxana Tarasenko Klyatchenko

Work place: National Technical University of Ukraine “Igor Sikorsky Kyiv Polytechnic Institute”, Kyiv, 03056, Ukraine

E-mail: o.tarasenko-kliatchenko@kpi.ua

Website: https://orcid.org/0000-0001-7200-0913

Research Interests:

Biography

Oxana Tarasenko Klyatchenko PhD, associate professor of Department of System Programming and Specialized Computer Systems of National Technical University of Ukraine “Igor Sikorsky Kyiv Polytechnic Institute”, Ukraine. Scientific interests: programming, computer logic, IoT, big data analytics, machine learning.

Author Articles
The Efficiency of FPGA-based Matrix Multiplication Accelerators for Neural Network Algorithms

By Yaroslav Klyatchenko Oxana Tarasenko Klyatchenko

DOI: https://doi.org/10.5815/ijem.2026.05.12, Pub. Date: 8 Oct. 2026

The contemporary challenges facing computer engineering have led to a focus on improving the efficiency of hardware and computing systems, as well as on optimising the operation of devices for effective data processing. In most modern systems of artificial intelligence, computer vision and digital signal processing, matrix multiplication is a basic operation. With the increasing resolution of sensors and the growing complexity of neural networks, classical general-purpose processors face the ‘von Neumann bottleneck’, where the data transfer rate between memory and the processor is limited, and the sequential execution of instructions does not allow the required real-time throughput to be achieved. The subject of this research is a System-on-a-Chip architecture that combines a dual-core processor based on the ARM architecture with programmable logic. This hybrid structure allows the most computationally intensive tasks to be offloaded to the hardware, whilst leaving control and the implementation of high-level interfaces to software. A neural network accelerator based on programmable logic devices offers advantages such as the capability for stream processing, which minimises the number of accesses to external memory, and the ability to utilise so-called mixed-precision computing. Furthermore, the accelerator’s efficiency is achieved through the use of multiple processing elements, enabling parallel computation of the neural network’s output channels. It has been demonstrated that implementing matrix operations on programmable logic enables parallelism of hundreds of operations per clock cycle, which is unattainable for systems based on general-purpose CPUs of the same class. The proposed implementations for a programmable logic-based accelerator have demonstrated speed-up compared to a system based on an ARM-architecture processor, as well as high adaptability to Deep Learning tasks.

[...] Read more.
Other Articles