Closed-loop Real-time Safety Critic as Guardrail in Robotics
Hämtar...
Ladda ner
Publicerad
Författare
Typ
Examensarbete för masterexamen
Master's Thesis
Master's Thesis
Modellbyggare
Tidskriftstitel
ISSN
Volymtitel
Utgivare
Sammanfattning
Robots are increasingly used in manufacturing, logistics, service, and manipulation
settings, where physical safety is a basic premise rather than an optional supplement.
While robot policies may be obtained through reinforcement learning, imitation
learning, optimal control, etc., the development normally involves training, simulation,
testing, or rollout data. Apart from learning task behavior, the interactive data can
also be used to extract knowledge about danger. To address this issue, this thesis
adapts the idea of safe reinforcement learning to a deployment setting in which the
task policy is already available and should remain frozen. Instead of retraining the
nominal policy with a new constrained objective, the work asks whether a learned
safety module can act as a last-step guardrail: before each action is executed, it
evaluates the nominal action, accepts it when risk is low, applies a small correction
when risk is moderate, or switches to a recovery action when collision risk is high.
Specifically, the proposed method is a closed-loop robot safety guardrail built around
a real-time dual-head safety critic. Mixed-behavior trajectories are first used to
construct hard-collision labels and continuous soft-risk labels; the critic is then
pre-trained to estimate hard-constraint violation risk and cumulative soft-risk cost.
During online execution, directional gating, candidate triggering, random shooting,
gradient-projection soft correction, and emergency Recovery Actor takeover are
combined to produce hierarchical and minimally invasive action correction.
The framework is evaluated on multiple task suites, including a PointMaze navigation
problem, a Franka Panda reach-obstacle task, and a Meta-World pick-place-wall
manipulation, where the proposed guardrail consistently reduces collision risk while
preserving task performance. To start with, collision rates prominently drop in all
scenarios. For Panda reach-obstacle, the average minimum clearance also improves.
On Meta-World, while the collision instances dramatically decrease, the overall task
success also rises. Additional ablation results show that soft correction and Recovery
address different risk stages and work best when combined.
Beskrivning
Ämne/nyckelord
safety in robotics, safe reinforcement learning, collision avoidance
