renderfy/runopsy-bench
Runopsy-Bench Twenty labelled agent traces for measuring failure-onset localization: given a run that went wrong, which step did it start going wrong at — not which step it stopped at. Produced for Runopsy, an open-source causal failure analysis engine for agent runs. pip install runopsy. Read this first: these traces are synthetic Every case here was generated, not recorded. They are single-fault traces written to exercise a specific failure mode, with the onset… See the full description on the dataset page: https://huggingface.co/datasets/renderfy/runopsy-bench.
013
Runopsy-Bench: 20 labelled traces for failure-onset localization
initial commit
