datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
varroa_mmdet_yolo_protocol_runsNOMOS-GEO-Audit-Protocol
NOMOS GEO Audit Protocol
A repeatable way to test what AI systems say about an organisation and whether the evidence supports it
GEO means Generative Engine Optimization. This six-language candidate protocol turns that discipline into an auditable process using the GEO-1000 method, canonical questions, truth packs, evidence requirements, scoring logic, correction steps and revalidation records.
Start reading: Open the English PDF · Choose one of six languages · Cite… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/NOMOS-GEO-Audit-Protocol.agentic-publication-protocol-dataset
APP compare-app benchmark
Paired reader conversations and blinded evaluations comparing an Agentic
Publication Protocol (APP) paper agent against a general repository-aware
agent, on 11 quantum-physics papers.
For each paper, a neutral reader asks the same scripted questions to both agents;
the two transcripts are anonymized and scored by a blinded evaluator on
accuracy, informativeness, grounding, and honesty (1-10).
Evaluator: Codex CLI, gpt-5.5, reasoning effort xhigh… See the full description on the dataset page: https://huggingface.co/datasets/phynics/agentic-publication-protocol-dataset.NOMOS-GBO-Audit-Protocol
NOMOS GBO Audit Protocol
Prove what an AI agent did, what authorised it, which evidence supports the judgement and whether it could be stopped.
Saying that an AI agent followed its instructions is not evidence. The NOMOS GBO Audit Protocol turns Generative Behavior Optimization (GBO) into a practical method for testing authority, tool use, evidence, delegation, stopping, recovery and human control.
Start reading: Open the English PDF · Choose one of six editions
Kaan… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/NOMOS-GBO-Audit-Protocol.protocol-bench
Protocol-Bench
15 published IEEE 802.11 and 3GPP procedures with ground-truth safety verdicts — and, where a
property fails, the shortest counterexample trace that proves it.
Most reasoning benchmarks accept an answer. This one asks for a proof: if a model says a protocol
is broken, it must supply a trace that starts at the initial state, moves only along real
transitions, and ends in a genuinely violating state. Traces are replayed mechanically. A
plausible-sounding trace that… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/protocol-bench.agentic-publication-protocol-dev-data
APP compare-app benchmark
Paired reader conversations and blinded evaluations comparing an Agentic
Publication Protocol (APP) paper agent against a general repository-aware
agent, on 11 public quantum-physics papers. This is the public-paper
subset reported in the APP paper's compare-app table.
For each paper, a neutral reader asks the same scripted questions to both agents;
the two transcripts are anonymized and scored by a blinded evaluator on
accuracy, informativeness… See the full description on the dataset page: https://huggingface.co/datasets/LionSR/agentic-publication-protocol-dev-data.
