CoolFace
Datasetpublic

Hachiman94/receipts-agent-claims

Receipts — Agent Claim Transcripts Every transcript from the Receipts benchmark — one row per trial, graded by a pytest exit code rather than by another model. 424 runs on claude-haiku-4-5, plus 6 pilot runs on gemini-2.5-flash via aider. All trials are committed. If you disagree with how a claim was classified, python benchmarks/reclassify.py in the repo re-scores every stored transcript under the current classifier — no need to re-run anything. What the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Hachiman94/receipts-agent-claims.

sourceHugging Facemitupdated 22d agoView on Hugging Face
0likes44downloads
2 commits on main
5dbc5b022d ago

Publish transcripts and dataset card

Hachiman94
c1dccea22d ago

initial commit

Hachiman94