Chiara Zavatta
Efficient and fast, low bitwidth post-training quantization.
Rel. Claudio Passerone, Pierpaolo Mori'. Politecnico di Torino, Corso di laurea magistrale in Ingegneria Informatica (Computer Engineering), 2026
|
Preview |
PDF (Tesi_di_laurea)
- Tesi
Licenza: Creative Commons Attribution Non-commercial No Derivatives. Download (6MB) | Preview |
Abstract
Edge AI (Artificial Intelligence) has become important in embedded systems because it enables AI algorithms to run directly on resource constrained devices. However, there are still several challenges to face when deploying neural networks on edge. For instance, while edge devices present strict constraints in terms of memory, energy and computation, state of the art DNNs have millions of parameters and require massive computational volumes. This makes the original model too slow for latency-critical applications and too large for memory-constrained hardware. To bridge this gap, model compression techniques, such as quantization, can be used to reduce memory demand and speed up inference, but often cause a severe degradation of the final accuracy, especially in low bit-width.
This thesis aims to optimize post-training quantization to reach a fast inference and reduced memory consumption while successfully preserving original task accuracy without the need for expensive or unfeasible model retraining
Relatori
Anno Accademico
Tipo di pubblicazione
Numero di pagine
Corso di laurea
Classe di laurea
Aziende collaboratrici
URI
![]() |
Modifica (riservato agli operatori) |
