SZLHOLDINGS/governed-agent-bench
governed-agent-bench v0 This dataset is the immutable public mirror of szl-holdings/a11oy@1b40fcbe0f65c1ad1e07776f83aa01abb067e864. It measures five governability axes: fail-closed behavior; non-increasing authority across delegation; false-success rejection; receipt completeness; and rollback discipline. Evidence labels Corpus: SAMPLE Scores: COMPUTED Receipt verification: STRUCTURE_ONLY Cryptographic verification: false The reference result proves that the… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/governed-agent-bench.
governed-agent-bench v0
This dataset is the immutable public mirror of `szl-holdings/a11oy@1b40fcbe0f65c1ad1e07776f83aa01abb067e864`.
It measures five governability axes:
- fail-closed behavior;
- non-increasing authority across delegation;
- false-success rejection;
- receipt completeness; and
- rollback discipline.
Evidence labels
- Corpus: SAMPLE
- Scores: COMPUTED
- Receipt verification: STRUCTURE_ONLY
- Cryptographic verification: false
The reference result proves that the evaluator and known-good fixture close their deterministic contract. It is not a model-quality or production claim. The public leaderboard contains zero eligible model submissions until an exact submission is evaluated and published with its receipt.
Reproduce
python score.py submissions/reference-conformance.jsonl --strictThe canonical source, schema, evaluator, reference submission, result, and publication manifest are all included in this dataset revision. The companion Space is <https://huggingface.co/spaces/SZLHOLDINGS/governed-agent-bench>.
