CoolFace
Datasetpublic

Hachiman94/receipts-agent-claims

Receipts — Agent Claim Transcripts Every transcript from the Receipts benchmark — one row per trial, graded by a pytest exit code rather than by another model. 424 runs on claude-haiku-4-5, plus 6 pilot runs on gemini-2.5-flash via aider. All trials are committed. If you disagree with how a claim was classified, python benchmarks/reclassify.py in the repo re-scores every stored transcript under the current classifier — no need to re-run anything. What the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Hachiman94/receipts-agent-claims.

sourceHugging Facemitupdated 22d agoView on Hugging Face
0likes44downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Hachiman94/receipts-agent-claims · CoolFace