Risk-Aware Safety in Multi-Agent Systems - Online CVaR Q-Learning, Q-Function–Multiplier Decoupling, and the Centralised Scaling Bottleneck
Hämtar...
Ladda ner
Publicerad
Författare
Typ
Examensarbete för masterexamen
Master's Thesis
Master's Thesis
Modellbyggare
Tidskriftstitel
ISSN
Volymtitel
Utgivare
Sammanfattning
Some plans look good on average and still go badly wrong. A taxi dispatcher routing
a fleet under shifting demand cannot judge its strategy by the typical shift alone:
the difference between strategies shows up on the worst shifts, when calls outpace
the fleet and passengers are left waiting. The standard framework for such problems
is the Markov decision process (MDP), and a standard way to control upper-tail
severity is Conditional Value-at-Risk (CVaR), which measures the average cost in
the worst few per cent of outcomes.
In practice the dispatcher has no advance model of how the shift will unfold. Each
shift differs from the last: an evening event reshapes demand, weather closes routes,
vehicles break down; the dispatcher must learn its policy from experience. The
problem is one of online reinforcement learning. Borkar and Jain (2014) gave the
canonical algorithm for this setting: a tabular Q-learner whose policy, tail threshold,
and constraint multiplier evolve together on three timescales. They proved that it
converges in the limit. Its finite-time behaviour, and its scaling to multi-agent
systems, remain open.
Here we show that the cost Q-function in this algorithm plays two roles — it scores
actions, and it drives the constraint multiplier — which coincide in the limit but
separate in finite time. We propose three changes that retain the asymptotic CVaR
target: we correct the literal cost-side update with a Bellman-consistent one; we
decouple the reward and cost Q-functions, running them as separate recursions that
recombine only at action selection; and we source the multiplier signal from a rolling
buffer of return samples rather than from the cost Q-function itself.
We then stress-test the resulting learner under centralisation, where one learner
acts over the joint state-action space of all agents. It clears the five-component
Karush–Kuhn–Tucker (KKT)-based operational gate at the calibrated single-agent
point and in the centralised lift with one and two agents; the wider single-agent grid
characterises the operating envelope. Under the same frozen centralised protocol
and per-seed training budget, it stalls on every seed at three agents. We trace the
stall to two structural quantities of the joint representation — the volume of its Q
function table and the variance of its cumulative cost — which are not targeted by
C1–C3 at fixed centralised representation, and motivate per-agent risk-contribution
allocation as the forward construction examined in the final chapter.
Beskrivning
Ämne/nyckelord
Conditional Value-at-Risk; online Q-learning; three-timescale stochastic approximation; Lagrangian relaxation; Q-function–multiplier decoupling; finite-time analysis; centralised multi-agent reinforcement learning; risk-contribution allocation.
