Risk-Aware Safety in Multi-Agent Systems - Online CVaR Q-Learning, Q-Function–Multiplier Decoupling, and the Centralised Scaling Bottleneck

dc.contributor.authorCheng, Xi
dc.contributor.departmentChalmers tekniska högskola / Institutionen för data och informationstekniksv
dc.contributor.departmentChalmers University of Technology / Department of Computer Science and Engineeringen
dc.contributor.examinerDamaschke, Peter
dc.contributor.supervisorGautier, Anna
dc.date.accessioned2026-06-29T13:54:42Z
dc.date.issued2026
dc.date.submitted
dc.description.abstractSome plans look good on average and still go badly wrong. A taxi dispatcher routing a fleet under shifting demand cannot judge its strategy by the typical shift alone: the difference between strategies shows up on the worst shifts, when calls outpace the fleet and passengers are left waiting. The standard framework for such problems is the Markov decision process (MDP), and a standard way to control upper-tail severity is Conditional Value-at-Risk (CVaR), which measures the average cost in the worst few per cent of outcomes. In practice the dispatcher has no advance model of how the shift will unfold. Each shift differs from the last: an evening event reshapes demand, weather closes routes, vehicles break down; the dispatcher must learn its policy from experience. The problem is one of online reinforcement learning. Borkar and Jain (2014) gave the canonical algorithm for this setting: a tabular Q-learner whose policy, tail threshold, and constraint multiplier evolve together on three timescales. They proved that it converges in the limit. Its finite-time behaviour, and its scaling to multi-agent systems, remain open. Here we show that the cost Q-function in this algorithm plays two roles — it scores actions, and it drives the constraint multiplier — which coincide in the limit but separate in finite time. We propose three changes that retain the asymptotic CVaR target: we correct the literal cost-side update with a Bellman-consistent one; we decouple the reward and cost Q-functions, running them as separate recursions that recombine only at action selection; and we source the multiplier signal from a rolling buffer of return samples rather than from the cost Q-function itself. We then stress-test the resulting learner under centralisation, where one learner acts over the joint state-action space of all agents. It clears the five-component Karush–Kuhn–Tucker (KKT)-based operational gate at the calibrated single-agent point and in the centralised lift with one and two agents; the wider single-agent grid characterises the operating envelope. Under the same frozen centralised protocol and per-seed training budget, it stalls on every seed at three agents. We trace the stall to two structural quantities of the joint representation — the volume of its Q function table and the variance of its cumulative cost — which are not targeted by C1–C3 at fixed centralised representation, and motivate per-agent risk-contribution allocation as the forward construction examined in the final chapter.
dc.identifier.coursecodeDATX05
dc.identifier.urihttps://hdl.handle.net/20.500.12380/311629
dc.language.isoeng
dc.setspec.uppsokTechnology
dc.subjectConditional Value-at-Risk; online Q-learning; three-timescale stochastic approximation; Lagrangian relaxation; Q-function–multiplier decoupling; finite-time analysis; centralised multi-agent reinforcement learning; risk-contribution allocation.
dc.titleRisk-Aware Safety in Multi-Agent Systems - Online CVaR Q-Learning, Q-Function–Multiplier Decoupling, and the Centralised Scaling Bottleneck
dc.type.degreeExamensarbete för masterexamensv
dc.type.degreeMaster's Thesisen
dc.type.uppsokH
local.programmeComplex adaptive systems (MPCAS), MSc

Ladda ner

Original bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
CSE 26-26 XC.pdf
Size:
1.61 MB
Format:
Adobe Portable Document Format

License bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
license.txt
Size:
2.35 KB
Format:
Item-specific license agreed upon to submission
Description: