corrigibility/corrigibility-intervention-lab
0
๐ Corrigibility Intervention Lab
A browser-based agent intervention evaluator by Corrigibility.
Simulate an agent before and after an authorized human intervention. Test whether the agent stops, updates goals, respects permission changes, follows constraints, waits for approval, and recovers safely.
Intervention types
- Stop / interrupt
- Goal update
- Permission revocation
- Approval gate
- Constraint change
- Tool removal
Evaluation dimensions
- Interruption compliance
- Goal flexibility
- Permission respect
- Constraint following
- Recovery quality
- Oversight escalation
- Trace transparency
Outputs
- Corrigibility profile
- Intervention timeline
- Pass / warning / fail signals
- Residual autonomy risk
- Recovery recommendation
- Machine-readable evaluation JSON
- Shareable scenario state
Important
This Space is a synthetic evaluation sandbox. It does not execute real agents, tools, payments, files, or external actions. Scores are heuristic and scenario-based, not a universal measure of corrigibility.
Privacy
The app runs entirely in the browser. Scenario inputs are not sent to a backend.
Tech
- Static HTML
- CSS
- Vanilla JavaScript
- No backend
- No API
- No external framework
Built for the Corrigibility organization on Hugging Face.
