Reinforcement Learning for Optimized Heat Plant Planning - Scheduling strategies in heat plant planning with long horizon forecasts and MILP optimality comparisons
| dc.contributor.author | Göransson Kjellmer, Karl | |
| dc.contributor.author | Lindqvist, Love | |
| dc.contributor.department | Chalmers tekniska högskola / Institutionen för data och informationsteknik | sv |
| dc.contributor.department | Chalmers University of Technology / Department of Computer Science and Engineering | en |
| dc.contributor.examiner | Haghir, Morteza | |
| dc.contributor.supervisor | Tamino Freitag, Kilian | |
| dc.date.accessioned | 2026-07-07T10:06:37Z | |
| dc.date.issued | 2026 | |
| dc.date.submitted | ||
| dc.description.abstract | District Heating Plants (DHPs) supply a significant share of heat for buildings in Sweden and present a non-trivial scheduling problem. Decisions about which production unit to run when and at which level must be balanced against time-varying electricity prices, heat demand, thermal storage dynamics, and unit-level operational constraints. Industry formulates this as a Mixed-Integer Linear Programming (MILP) problem, but MILP scales poorly with system size and horizon length. This thesis investigates whether Deep Reinforcement Learning (DRL), trained with week-ahead price and demand forecasts, can produce schedules competitive with a MILP baseline on the same problem instance. Two simulated DHP configurations of differing complexity are introduced, one corresponding to a household and one to a small town. Tested with a DRL Proximal Policy Optimization (PPO) agent that uses a convolutional encoder for forecast time series. Ablations on the action distribution (Gaussian, Beta, Kumaraswamy), the forecast encoder, the mechanism for enforcing demand satisfaction, and a transfer-learning study across plant configurations are presented. Against a rolling-horizon MILP on identical validation data, the trained agents learned a promising strategy reaching within 8.1% of MILP operating cost on the simpler configuration and effectively match it on the more complex one. However, the resulting schedules exhibit high-frequency unit switching and under-use of accumulator capacity which need to be addressed, still highlighting the potential of Reinforcement Learning (RL) in this domain. | |
| dc.identifier.uri | https://hdl.handle.net/20.500.12380/311906 | |
| dc.language.iso | eng | |
| dc.setspec.uppsok | Technology | |
| dc.subject | Reinforcement Learning, Proximal Policy Optimization, District Heating, Mixed-Integer Linear Programming, Scheduling, Distributions, Convolutional Neural Network, Deep Reinforcement Learning, Long Horizon | |
| dc.title | Reinforcement Learning for Optimized Heat Plant Planning - Scheduling strategies in heat plant planning with long horizon forecasts and MILP optimality comparisons | |
| dc.type.degree | Examensarbete för masterexamen | sv |
| dc.type.degree | Master's Thesis | en |
| dc.type.uppsok | H | |
| local.programme | Computer science -algorithms, languages and logic (MPALG), MSc |
