vedant33/supersede-demo
0
Supersede Supersession Demo
An interactive look at supersession: across a multi-session chat a fact changes, and the agent must answer with the current value, not a superseded one. Browse real episodes from the Supersede RL environment and see the current answer vs. the superseded traps.
GRPO training on this environment lifts held-out supersession accuracy on LongMemEval (knowledge-update, oracle) from 9.0% โ 16.7% (Qwen2.5-3B-Instruct).
