Accuracy and Hardware Cost Analysis of Multi-Format Floating-Point Arithmetic Generated by FloPoCo on FPGA
| dc.contributor.author | Sun, Boyue | |
| dc.contributor.author | Xia, Zhili | |
| dc.contributor.department | Chalmers tekniska högskola / Institutionen för mikroteknologi och nanovetenskap (MC2) | sv |
| dc.contributor.department | Chalmers University of Technology / Department of Microtechnology and Nanoscience (MC2) | en |
| dc.contributor.examiner | Peterson, Lena | |
| dc.contributor.supervisor | Larsson-Edefors, Per | |
| dc.date.accessioned | 2026-07-01T09:30:01Z | |
| dc.date.issued | 2026 | |
| dc.date.submitted | ||
| dc.description.abstract | Modern artificial intelligence (AI) and digital signal processing (DSP) workloads are highly data-driven and computationally demanding, relying heavily on massive multiply-accumulate (MAC) operations. While standard IEEE 754 floating-point arithmetic provides a vast dynamic range, its strict compliance requirements, such as subnormal handling and exact rounding, incur significant hardware overhead. To address this, this thesis evaluates the accuracy and hardware cost of multi format floating-point arithmetic generated by the FloPoCo framework on field programmable gate arrays (FPGAs). We employ a hardware-software co-simulation methodology, combining Xilinx Vivado for power, performance, and area assessment with a Python-based error evaluation engine using Gaussian distributed test vectors to emulate AI workloads. Our results demonstrate that FloPoCo’s custom Nfloat format, which eliminates subnormal support and utilizes a dedicated exception field, significantly reduces pipeline depth, look-up table (LUT) consumption, and dynamic power compared to IEEE 754 implementations across all tested bit-widths. Furthermore, a comparative analysis between a unified-precision Nfloat MAC and an IEEE fused multiply-add (FMA) reveals that the NFloat MAC achieves over 50% power and area savings while maintaining identical algorithmic fidelity at low-to-medium precisions. Finally, we investigate the performance of mixed-precision MAC architectures in deep accumulation chains with lengths up to 5120 accumulation steps. The results show that mixed-precision computation effectively maintains a stable relative error near 0.001%. FloPoCo and the Nfloat format present a efficient and customizable alter native for FPGA-based high-performance computing. | |
| dc.identifier.coursecode | MCCX04 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.12380/311732 | |
| dc.language.iso | eng | |
| dc.setspec.uppsok | PhysicsChemistryMaths | |
| dc.subject | FPGA,Floating-Point Arithmetic, FloPoCo, Nfloat, Multiply-Accumulate, Mixed-Precision, Hardware-Software Co-Simulation | |
| dc.title | Accuracy and Hardware Cost Analysis of Multi-Format Floating-Point Arithmetic Generated by FloPoCo on FPGA | |
| dc.type.degree | Examensarbete för masterexamen | sv |
| dc.type.degree | Master's Thesis | en |
| dc.type.uppsok | H | |
| local.programme | Embedded electronic system design (MPEES), MSc |
Ladda ner
Original bundle
1 - 1 av 1
Hämtar...
- Namn:
- Accuracy and Hardware Cost Analysis of Multi-Format Floating-Point Arithmetic Generated by FloPoCo on FPGA-final.pdf
- Size:
- 864.66 KB
- Format:
- Adobe Portable Document Format
License bundle
1 - 1 av 1
Hämtar...
- Namn:
- license.txt
- Size:
- 2.35 KB
- Format:
- Item-specific license agreed upon to submission
- Description:
