Accuracy and Hardware Cost Analysis of Multi-Format Floating-Point Arithmetic Generated by FloPoCo on FPGA

dc.contributor.authorSun, Boyue
dc.contributor.authorXia, Zhili
dc.contributor.departmentChalmers tekniska högskola / Institutionen för mikroteknologi och nanovetenskap (MC2)sv
dc.contributor.departmentChalmers University of Technology / Department of Microtechnology and Nanoscience (MC2)en
dc.contributor.examinerPeterson, Lena
dc.contributor.supervisorLarsson-Edefors, Per
dc.date.accessioned2026-07-01T09:30:01Z
dc.date.issued2026
dc.date.submitted
dc.description.abstractModern artificial intelligence (AI) and digital signal processing (DSP) workloads are highly data-driven and computationally demanding, relying heavily on massive multiply-accumulate (MAC) operations. While standard IEEE 754 floating-point arithmetic provides a vast dynamic range, its strict compliance requirements, such as subnormal handling and exact rounding, incur significant hardware overhead. To address this, this thesis evaluates the accuracy and hardware cost of multi format floating-point arithmetic generated by the FloPoCo framework on field programmable gate arrays (FPGAs). We employ a hardware-software co-simulation methodology, combining Xilinx Vivado for power, performance, and area assessment with a Python-based error evaluation engine using Gaussian distributed test vectors to emulate AI workloads. Our results demonstrate that FloPoCo’s custom Nfloat format, which eliminates subnormal support and utilizes a dedicated exception field, significantly reduces pipeline depth, look-up table (LUT) consumption, and dynamic power compared to IEEE 754 implementations across all tested bit-widths. Furthermore, a comparative analysis between a unified-precision Nfloat MAC and an IEEE fused multiply-add (FMA) reveals that the NFloat MAC achieves over 50% power and area savings while maintaining identical algorithmic fidelity at low-to-medium precisions. Finally, we investigate the performance of mixed-precision MAC architectures in deep accumulation chains with lengths up to 5120 accumulation steps. The results show that mixed-precision computation effectively maintains a stable relative error near 0.001%. FloPoCo and the Nfloat format present a efficient and customizable alter native for FPGA-based high-performance computing.
dc.identifier.coursecodeMCCX04
dc.identifier.urihttps://hdl.handle.net/20.500.12380/311732
dc.language.isoeng
dc.setspec.uppsokPhysicsChemistryMaths
dc.subjectFPGA,Floating-Point Arithmetic, FloPoCo, Nfloat, Multiply-Accumulate, Mixed-Precision, Hardware-Software Co-Simulation
dc.titleAccuracy and Hardware Cost Analysis of Multi-Format Floating-Point Arithmetic Generated by FloPoCo on FPGA
dc.type.degreeExamensarbete för masterexamensv
dc.type.degreeMaster's Thesisen
dc.type.uppsokH
local.programmeEmbedded electronic system design (MPEES), MSc

Ladda ner

Original bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
Accuracy and Hardware Cost Analysis of Multi-Format Floating-Point Arithmetic Generated by FloPoCo on FPGA-final.pdf
Size:
864.66 KB
Format:
Adobe Portable Document Format

License bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
license.txt
Size:
2.35 KB
Format:
Item-specific license agreed upon to submission
Description: