Closed-loop Real-time Safety Critic as Guardrail in Robotics

Hämtar...
Bild (thumbnail)

Publicerad

Typ

Examensarbete för masterexamen
Master's Thesis

Modellbyggare

Tidskriftstitel

ISSN

Volymtitel

Utgivare

Sammanfattning

Robots are increasingly used in manufacturing, logistics, service, and manipulation settings, where physical safety is a basic premise rather than an optional supplement. While robot policies may be obtained through reinforcement learning, imitation learning, optimal control, etc., the development normally involves training, simulation, testing, or rollout data. Apart from learning task behavior, the interactive data can also be used to extract knowledge about danger. To address this issue, this thesis adapts the idea of safe reinforcement learning to a deployment setting in which the task policy is already available and should remain frozen. Instead of retraining the nominal policy with a new constrained objective, the work asks whether a learned safety module can act as a last-step guardrail: before each action is executed, it evaluates the nominal action, accepts it when risk is low, applies a small correction when risk is moderate, or switches to a recovery action when collision risk is high. Specifically, the proposed method is a closed-loop robot safety guardrail built around a real-time dual-head safety critic. Mixed-behavior trajectories are first used to construct hard-collision labels and continuous soft-risk labels; the critic is then pre-trained to estimate hard-constraint violation risk and cumulative soft-risk cost. During online execution, directional gating, candidate triggering, random shooting, gradient-projection soft correction, and emergency Recovery Actor takeover are combined to produce hierarchical and minimally invasive action correction. The framework is evaluated on multiple task suites, including a PointMaze navigation problem, a Franka Panda reach-obstacle task, and a Meta-World pick-place-wall manipulation, where the proposed guardrail consistently reduces collision risk while preserving task performance. To start with, collision rates prominently drop in all scenarios. For Panda reach-obstacle, the average minimum clearance also improves. On Meta-World, while the collision instances dramatically decrease, the overall task success also rises. Additional ablation results show that soft correction and Recovery address different risk stages and work best when combined.

Beskrivning

Ämne/nyckelord

safety in robotics, safe reinforcement learning, collision avoidance

Citation

Arkitekt (konstruktör)

Geografisk plats

Byggnad (typ)

Byggår

Modelltyp

Skala

Teknik / material

Index

Endorsement

Review

Supplemented By

Referenced By