datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
governed-agent-bench
governed-agent-bench v0
This dataset is the immutable public mirror of
szl-holdings/a11oy@1b40fcbe0f65c1ad1e07776f83aa01abb067e864.
It measures five governability axes:
fail-closed behavior;
non-increasing authority across delegation;
false-success rejection;
receipt completeness; and
rollback discipline.
Evidence labels
Corpus: SAMPLE
Scores: COMPUTED
Receipt verification: STRUCTURE_ONLY
Cryptographic verification: false
The reference result proves that the… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/governed-agent-bench.agentforge-governed-rag-evals
AgentForge Governed RAG Evaluation Set
Small, reproducible evaluation fixtures from AgentForge Enterprise.
golden.jsonl: expected evidence/approval contract cases.
knowledge_base.jsonl: the deterministic knowledge base used by the reference retriever.
Source revision: bf9eda819ae2b714b87400ee4188fa4e201cf275.
This dataset measures engineering contracts, not general LLM intelligence or enterprise production accuracy.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/singhankit491/agentforge-governed-rag-evals.
