CoolFace
Datasetpublic

SZLHOLDINGS/governed-agent-bench

governed-agent-bench v0 This dataset is the immutable public mirror of szl-holdings/a11oy@1b40fcbe0f65c1ad1e07776f83aa01abb067e864. It measures five governability axes: fail-closed behavior; non-increasing authority across delegation; false-success rejection; receipt completeness; and rollback discipline. Evidence labels Corpus: SAMPLE Scores: COMPUTED Receipt verification: STRUCTURE_ONLY Cryptographic verification: false The reference result proves that the… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/governed-agent-bench.

sourceHugging Faceapache-2.0updated 25d agoView on Hugging Face
0likes355downloads
Dataset Card

governed-agent-bench v0

This dataset is the immutable public mirror of `szl-holdings/a11oy@1b40fcbe0f65c1ad1e07776f83aa01abb067e864`.

It measures five governability axes:

  1. 1.fail-closed behavior;
  2. 2.non-increasing authority across delegation;
  3. 3.false-success rejection;
  4. 4.receipt completeness; and
  5. 5.rollback discipline.

Evidence labels

  • —Corpus: SAMPLE
  • —Scores: COMPUTED
  • —Receipt verification: STRUCTURE_ONLY
  • —Cryptographic verification: false

The reference result proves that the evaluator and known-good fixture close their deterministic contract. It is not a model-quality or production claim. The public leaderboard contains zero eligible model submissions until an exact submission is evaluated and published with its receipt.

Reproduce

bash
python score.py submissions/reference-conformance.jsonl --strict

The canonical source, schema, evaluator, reference submission, result, and publication manifest are all included in this dataset revision. The companion Space is <https://huggingface.co/spaces/SZLHOLDINGS/governed-agent-bench>.