article
Neural network quantization is a key enabler for deploying deep learning models on field-programmable gate arrays (FPGAs), as it significantly reduces memory footprint, computational complexity, and energy consumption while preserving competitive inference accuracy. By exploiting customized datapaths and fine-grained parallelism, FPGAs are particularly well suited for low-precision arithmetic. Nevertheless, the rapid expansion of quantization techniques, FPGA platforms, and deployment frameworks has given rise to a complex design space in which selecting suitable solutions under accuracy, performance, and resource constraints remains challenging. This paper introduces a unified taxonomy and presents a focused survey of neural network quantization for FPGA deployment. Existing approaches are organized along key design dimensions, including quantization stage, numerical precision, granularity, hardware awareness, and deployment abstraction level. Primary techniques, such as post-training quantization and quantizationaware training, are reviewed alongside advanced strategies encompassing mixed-precision, extreme low-bit, and hardwareaware quantization. In addition, representative FPGA platforms and deployment frameworks are analyzed, and major application domains benefiting from quantized FPGA-based inference are summarized. Finally, open challenges are discussed and future research directions toward more automated, portable, and efficient FPGA-based deep learning systems are outlined.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/iraset68627.2026.11538436
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.