datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
receipts-agent-claims
Receipts — Agent Claim Transcripts
Every transcript from the Receipts benchmark — one row per trial, graded by a pytest exit code rather than by another model.
424 runs on claude-haiku-4-5, plus 6 pilot runs on gemini-2.5-flash via aider. All trials are committed. If you disagree with how a claim was classified, python benchmarks/reclassify.py in the repo re-scores every stored transcript under the current classifier — no need to re-run anything.
What the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Hachiman94/receipts-agent-claims.chrononoise-claims-fr
ChronoNoise-Claims-FR
ChronoNoise-Claims-FR is a silver dataset for studying whether LLM-generated historical claims are supported by noisy OCR text, post-corrected text, and document metadata.
The dataset is designed around a common historical NLP failure chain:
historical OCR noise → fluent LLM interpretation → plausible but unsupported historical claim
It is derived from a ChronoCorrect-Europeana-style dataset built from historical French newspaper OCR. Each record contains… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/chrononoise-claims-fr.grpo-oumi-synthetic-document-claims
Dataset Card for GRPO Oumi ANLI Subset
Dataset
This dataset is a reformatted version of the oumi-ai/oumi-synthetic-document-claims dataset, specifically structured for use with the GRPO trainer.
You can find more detailed information about the original dataset at the provided link.
Link: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-document-claims
Dataset Structure
The dataset consists of a list of dictionaries, where each dictionary represents a… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-synthetic-document-claims.grpo-oumi-synthetic-claims
Dataset Card for GRPO Oumi ANLI Subset
Dataset
This dataset is a reformatted version of the TEEN-D/grpo-oumi-anli-subset dataset, specifically structured for use with the GRPO trainer.
You can find more detailed information about the original dataset at the provided link.
Link: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-claims
Dataset Structure
The dataset consists of a list of dictionaries, where each dictionary represents a single data instance… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-synthetic-claims.
