Giovanni Pellegrino
Pruning-Aware Analysis of Accumulator Bit-Width in TFHE-Based CNN Inference.
Rel. Valentino Peluso, Andrea Calimera. Politecnico di Torino, Corso di laurea magistrale in Data Science And Engineering, 2026
|
Preview |
PDF (Tesi_di_laurea)
- Tesi
Licenza: Creative Commons Attribution Non-commercial No Derivatives. Download (3MB) | Preview |
Abstract
This thesis investigates how quantization and pruning can reduce the computational cost of convolutional neural network inference under Fully Homomorphic Encryption, with specific focus on TFHE-based execution. Although FHE enables privacy-preserving inference directly on encrypted data, its practical adoption is limited by the high latency of homomorphic operations, especially programmable bootstrapping. A central bottleneck is the accumulator bit-width required by convolutional layers, since larger intermediate sums increase ciphertext parameters, noise growth, and execution time. The work first reviews FHE, TFHE, quantized neural networks, and pruning strategies, then proposes an experimental methodology to analyze their interaction. A controlled layer-wise study on ResNet-18 evaluates how sparsity and quantization affect accumulator requirements and inference latency in representative convolutional layers.
This is complemented by an end-to-end exploration on ResNet-8 over CIFAR-10, where quantization-aware training, iterative pruning, and accuracy constraints are jointly considered
Relatori
Anno Accademico
Tipo di pubblicazione
Numero di pagine
Corso di laurea
Classe di laurea
Aziende collaboratrici
URI
![]() |
Modifica (riservato agli operatori) |
