datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
receipts-agent-claims
Receipts — Agent Claim Transcripts
Every transcript from the Receipts benchmark — one row per trial, graded by a pytest exit code rather than by another model.
424 runs on claude-haiku-4-5, plus 6 pilot runs on gemini-2.5-flash via aider. All trials are committed. If you disagree with how a claim was classified, python benchmarks/reclassify.py in the repo re-scores every stored transcript under the current classifier — no need to re-run anything.
What the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Hachiman94/receipts-agent-claims.answers-with-receipts
Answers with Receipts
26 real customer-support questions, each answered by an autonomous AI agent that paid its own money to compete, and each answer approved by the business that asked the question. Every row carries the on-chain transaction that paid the agent.
The preference label in this dataset is backed by a payment, not a click.
Why this is unusual
Most human-feedback datasets label a preference with an annotator's click. A click is cheap and reversible… See the full description on the dataset page: https://huggingface.co/datasets/deskcrew/answers-with-receipts.moralitylab-public-receipts
MoralityLab Public Research Snapshot
Dataset Summary
Public-safe snapshot of MoralityLab research artifacts used for reproducible reporting:
run manifests,
adapter/TRM indexes,
selected papers and docs,
dashboard-facing summary JSON.
This dataset intentionally excludes secrets, private credentials, and restricted raw traces.
Intended Uses
Public grant/research context.
Dashboard demo payloads for Harness/Gym.
Lightweight reproducibility receipts.… See the full description on the dataset page: https://huggingface.co/datasets/AlephFunk/moralitylab-public-receipts.
