Kandawalage Rukshike Stephano Perera
Development of an efficient matrix multiplication unit using High-Level Synthesis to accelerate attention-based neural network workloads on FPGA.
Rel. Luciano Lavagno, Roberto Bosio. Politecnico di Torino, Corso di laurea magistrale in Ingegneria Elettronica (Electronic Engineering), 2026
|
Preview |
PDF (Tesi_di_laurea)
- Tesi
Licenza: Creative Commons Attribution Non-commercial No Derivatives. Download (874kB) | Preview |
Abstract
This thesis aims to develop an FPGA-based framework for accelerating matrix multiplication in attention mechanisms using High-Level Synthesis and the NN2FPGA workflow. The need for this work arises from the growing computational demands of lightweight object detection models such as YOLOv10n, where low latency and efficient use of hardware resources are essential. Existing software-oriented implementations are often not well suited to these constraints, which motivates the exploration of a hardware-based solution. To address this challenge, the ONNX network is quantized using symmetric power of two integer quantization to reduce computational complexity and improve hardware efficiency. A streaming-based computation approach is then adopted to organize the accelerator dataflow and support efficient execution on FPGA.
The proposed framework is implemented and evaluated through synthesis, system integration and hardware testing on an FPGA platform
Relatori
Anno Accademico
Tipo di pubblicazione
Numero di pagine
Corso di laurea
Classe di laurea
Aziende collaboratrici
URI
![]() |
Modifica (riservato agli operatori) |
