Closed-loop Real-time Safety Critic as Guardrail in Robotics

dc.contributor.authorChen, Xingyu
dc.contributor.authorNan, Shipeng
dc.contributor.departmentChalmers tekniska högskola / Institutionen för data och informationstekniksv
dc.contributor.departmentChalmers University of Technology / Department of Computer Science and Engineeringen
dc.contributor.examinerAli-Eldin Hassan, Ahmed
dc.contributor.supervisorZhang, Ze
dc.date.accessioned2026-08-25T12:11:21Z
dc.date.issued2026
dc.date.submitted
dc.description.abstractRobots are increasingly used in manufacturing, logistics, service, and manipulation settings, where physical safety is a basic premise rather than an optional supplement. While robot policies may be obtained through reinforcement learning, imitation learning, optimal control, etc., the development normally involves training, simulation, testing, or rollout data. Apart from learning task behavior, the interactive data can also be used to extract knowledge about danger. To address this issue, this thesis adapts the idea of safe reinforcement learning to a deployment setting in which the task policy is already available and should remain frozen. Instead of retraining the nominal policy with a new constrained objective, the work asks whether a learned safety module can act as a last-step guardrail: before each action is executed, it evaluates the nominal action, accepts it when risk is low, applies a small correction when risk is moderate, or switches to a recovery action when collision risk is high. Specifically, the proposed method is a closed-loop robot safety guardrail built around a real-time dual-head safety critic. Mixed-behavior trajectories are first used to construct hard-collision labels and continuous soft-risk labels; the critic is then pre-trained to estimate hard-constraint violation risk and cumulative soft-risk cost. During online execution, directional gating, candidate triggering, random shooting, gradient-projection soft correction, and emergency Recovery Actor takeover are combined to produce hierarchical and minimally invasive action correction. The framework is evaluated on multiple task suites, including a PointMaze navigation problem, a Franka Panda reach-obstacle task, and a Meta-World pick-place-wall manipulation, where the proposed guardrail consistently reduces collision risk while preserving task performance. To start with, collision rates prominently drop in all scenarios. For Panda reach-obstacle, the average minimum clearance also improves. On Meta-World, while the collision instances dramatically decrease, the overall task success also rises. Additional ablation results show that soft correction and Recovery address different risk stages and work best when combined.
dc.identifier.coursecodeDATX05
dc.identifier.urihttps://hdl.handle.net/20.500.12380/312262
dc.language.isoeng
dc.setspec.uppsokTechnology
dc.subjectsafety in robotics, safe reinforcement learning, collision avoidance
dc.titleClosed-loop Real-time Safety Critic as Guardrail in Robotics
dc.type.degreeExamensarbete för masterexamensv
dc.type.degreeMaster's Thesisen
dc.type.uppsokH
local.programmeData science and AI (MPDSC), MSc

Ladda ner

Original bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
CSE 26-161 XC SN.pdf
Size:
4.27 MB
Format:
Adobe Portable Document Format

License bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
license.txt
Size:
2.35 KB
Format:
Item-specific license agreed upon to submission
Description: