Reinforcement Learning for Optimized Heat Plant Planning - Scheduling strategies in heat plant planning with long horizon forecasts and MILP optimality comparisons

dc.contributor.authorGöransson Kjellmer, Karl
dc.contributor.authorLindqvist, Love
dc.contributor.departmentChalmers tekniska högskola / Institutionen för data och informationstekniksv
dc.contributor.departmentChalmers University of Technology / Department of Computer Science and Engineeringen
dc.contributor.examinerHaghir, Morteza
dc.contributor.supervisorTamino Freitag, Kilian
dc.date.accessioned2026-07-07T10:06:37Z
dc.date.issued2026
dc.date.submitted
dc.description.abstractDistrict Heating Plants (DHPs) supply a significant share of heat for buildings in Sweden and present a non-trivial scheduling problem. Decisions about which production unit to run when and at which level must be balanced against time-varying electricity prices, heat demand, thermal storage dynamics, and unit-level operational constraints. Industry formulates this as a Mixed-Integer Linear Programming (MILP) problem, but MILP scales poorly with system size and horizon length. This thesis investigates whether Deep Reinforcement Learning (DRL), trained with week-ahead price and demand forecasts, can produce schedules competitive with a MILP baseline on the same problem instance. Two simulated DHP configurations of differing complexity are introduced, one corresponding to a household and one to a small town. Tested with a DRL Proximal Policy Optimization (PPO) agent that uses a convolutional encoder for forecast time series. Ablations on the action distribution (Gaussian, Beta, Kumaraswamy), the forecast encoder, the mechanism for enforcing demand satisfaction, and a transfer-learning study across plant configurations are presented. Against a rolling-horizon MILP on identical validation data, the trained agents learned a promising strategy reaching within 8.1% of MILP operating cost on the simpler configuration and effectively match it on the more complex one. However, the resulting schedules exhibit high-frequency unit switching and under-use of accumulator capacity which need to be addressed, still highlighting the potential of Reinforcement Learning (RL) in this domain.
dc.identifier.urihttps://hdl.handle.net/20.500.12380/311906
dc.language.isoeng
dc.setspec.uppsokTechnology
dc.subjectReinforcement Learning, Proximal Policy Optimization, District Heating, Mixed-Integer Linear Programming, Scheduling, Distributions, Convolutional Neural Network, Deep Reinforcement Learning, Long Horizon
dc.titleReinforcement Learning for Optimized Heat Plant Planning - Scheduling strategies in heat plant planning with long horizon forecasts and MILP optimality comparisons
dc.type.degreeExamensarbete för masterexamensv
dc.type.degreeMaster's Thesisen
dc.type.uppsokH
local.programmeComputer science -algorithms, languages and logic (MPALG), MSc

Ladda ner

Original bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
CSE 26-61 LL KGK.pdf
Size:
6.35 MB
Format:
Adobe Portable Document Format

License bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
license.txt
Size:
2.35 KB
Format:
Item-specific license agreed upon to submission
Description: