Cost-Latency Benchmarking for Time-Series Forecasting - Hardware-Aware Evaluation of Deep Learning Architectures for Financial Planning
| dc.contributor.author | Asplund, Gustaf | |
| dc.contributor.author | Bjerhem Aronsson, Felix | |
| dc.contributor.department | Chalmers tekniska högskola / Institutionen för data och informationsteknik | sv |
| dc.contributor.department | Chalmers University of Technology / Department of Computer Science and Engineering | en |
| dc.contributor.examiner | Ali-Eldin Hassan, Ahmed | |
| dc.contributor.supervisor | Ali-Eldin Hassan, Ahmed | |
| dc.date.accessioned | 2026-09-25T14:14:05Z | |
| dc.date.issued | 2026 | |
| dc.date.submitted | ||
| dc.description.abstract | Enterprise financial planning is increasingly moving from statistical forecasting toward deep learning, which captures more complex patterns at the product level but raises the cost of serving forecasts. The hardware for these workloads is often chosen through heuristics rather than measurement, and existing serving research has focused on computer vision and language rather than time-series forecasting. This thesis develops a workload-aware benchmarking framework that characterizes the cost-latency trade-offs of deep learning forecasting architectures across commodity cloud hardware. Six forecasting models spanning distinct computational classes were benchmarked on eleven Azure instances, examining how the operational intensity of each architecture sits against the roofline limits of the hardware, how far cost-latency behavior on real enterprise resource planning data diverges from simpler synthetic data, and how a Pareto analysis can guide the choice of a hardware and model pair under a given latency constraint. Operational intensity stayed within a narrow band across the tested models, yet the point at which a workload becomes compute- or memory-bound shifted with the hardware, so the same model could be bound differently from one machine to the next. The cost-latency outcome thus depends on the model and hardware as a pair rather than on the architecture alone, and this shows in the provisioning results. Neither the largest CPU nor the newest accelerator reliably improved serving performance and an older, lower-tier GPU instance offered the best overall balance for the workloads tested. The framework itself, rather than any single figure, is the more transferable result, and the comparison with synthetic data suggests it can stand in as a cheap first pass for narrowing the hardware search before real-data runs settle a final deployment. | |
| dc.identifier.coursecode | DATX05 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.12380/312559 | |
| dc.language.iso | eng | |
| dc.setspec.uppsok | Technology | |
| dc.subject | Time-series forecasting, deep learning, benchmarking, roofline model, cost-latency trade-off, cloud computing, inference serving, hardware selection, oper ational intensity, enterprise financial planning | |
| dc.title | Cost-Latency Benchmarking for Time-Series Forecasting - Hardware-Aware Evaluation of Deep Learning Architectures for Financial Planning | |
| dc.type.degree | Examensarbete för masterexamen | sv |
| dc.type.degree | Master's Thesis | en |
| dc.type.uppsok | H | |
| local.programme | High-performance computer systems (MPHPC), MSc |
