datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cms-desynpuf-insurance-claims
CMS DE-SynPUF Insurance Claims (Rendered EOB Documents)
Pipeline-ready insurance_claim evaluation corpus derived from the CMS
2008-2010 Data Entrepreneurs' Synthetic Public Use File (DE-SynPUF), Sample 1.
One row = one Medicare claim rendered as a plain-text EOB-style document with
ground-truth extraction fields aligned to the llm-mailroom
InsuranceClaimExtraction schema.
Provenance
Source: CMS DE-SynPUF Sample 1 (fully synthetic; "very limited inferential… See the full description on the dataset page: https://huggingface.co/datasets/Lucius-Morningstar/cms-desynpuf-insurance-claims.small-claims-court-dollar-limits-by-state
Small Claims Court Dollar Limits by State
Canonical, always-current version: https://referencesource.org/small-claims-court-dollar-limits-by-state/
Machine-readable: https://referencesource.org/small-claims-court-dollar-limits-by-state/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-26
Stale after: 2027-08-26 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 25
Every state sets its own maximum… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/small-claims-court-dollar-limits-by-state.receipts-agent-claims
Receipts — Agent Claim Transcripts
Every transcript from the Receipts benchmark — one row per trial, graded by a pytest exit code rather than by another model.
424 runs on claude-haiku-4-5, plus 6 pilot runs on gemini-2.5-flash via aider. All trials are committed. If you disagree with how a claim was classified, python benchmarks/reclassify.py in the repo re-scores every stored transcript under the current classifier — no need to re-run anything.
What the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Hachiman94/receipts-agent-claims.ai-job-claimsClaims about AI's impact on jobs from experts quoted in the news articles collected in Florent Daudens' AI Job News dataset
verified-step-claims-1k-preview
Verified Step Claims 1K (Preview)
Mathematically verified step-level claims with witness evidence for process supervision research.
This is a 1K preview sampled from a larger in-progress corpus built from thousands of verified math solvers, covering Korean CSAT (수능), national mock exams, and advanced textbook problems across calculus, probability, sequences, and more.
What Makes This Different
PRM800K (OpenAI)
This Dataset
Label source
Human annotators… See the full description on the dataset page: https://huggingface.co/datasets/limmak/verified-step-claims-1k-preview.medical-shorts-silver-claims-benchmark
Medical Shorts Silver Claims Benchmark
This dataset contains silver labels for evaluating medical claim detection and verification in short-form YouTube videos.
The dataset does not include video files, audio files, thumbnails, or keyframes.
It only includes YouTube video IDs, metadata, normalized medical claims, silver labels, and evidence source references.
Dataset Size
Videos: 236
Silver claims: 470
Visual-dependent claims: 22
Videos with medical claims: 205
Videos… See the full description on the dataset page: https://huggingface.co/datasets/ArtemkaT08/medical-shorts-silver-claims-benchmark.facebook-anli-claims
Facebook ANLI Claims Dataset
This dataset is a reformatted version of the facebook/anli dataset, reorganized by their original labels.
The original dataset's splits have been combined and separated into three distinct dataset configurations based on their labels: entailment, neutral, and contradiction.
For each configuration:
Input: Premise text
Output: A JSON list of hypotheses associated with the premise.
grpo-oumi-synthetic-document-claims
Dataset Card for GRPO Oumi ANLI Subset
Dataset
This dataset is a reformatted version of the oumi-ai/oumi-synthetic-document-claims dataset, specifically structured for use with the GRPO trainer.
You can find more detailed information about the original dataset at the provided link.
Link: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-document-claims
Dataset Structure
The dataset consists of a list of dictionaries, where each dictionary represents a… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-synthetic-document-claims.vietnamese-fact-checking-claims
Vietnamese Fact-Checking Claims
Generated claim-verification data derived from the Vietnamese Evidence Corpus.
Each article contains claims labeled as supported, refuted, or not having enough
information, together with evidence and a short rationale.
Statistics
12,238 source articles
73,454 generated claims
24,476 SUPPORTED claims
24,502 REFUTED claims
24,476 NOT_ENOUGH_INFO claims
Main fields
Article: id, date_iso, full_text, claims
Claim: claim… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-fact-checking-claims.echr-claimsgrpo-oumi-synthetic-claims
Dataset Card for GRPO Oumi ANLI Subset
Dataset
This dataset is a reformatted version of the TEEN-D/grpo-oumi-anli-subset dataset, specifically structured for use with the GRPO trainer.
You can find more detailed information about the original dataset at the provided link.
Link: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-claims
Dataset Structure
The dataset consists of a list of dictionaries, where each dictionary represents a single data instance… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-synthetic-claims.standing-desk-claims-evidence-2026
DeskDeploy 2026 Standing Desk Health Claims vs Evidence Audit
Brand-by-brand audit of 30 standing desk products mapping each marketed health claim to peer-reviewed evidence with Cochrane/PubMed evidence grades and verdicts.
License: CC-BY 4.0
DOI: 10.5281/zenodo.20632751
Source study: https://deskdeploy.com/research/standing-desk-claims-evidence-2026/
Author: Vincent Wesley Couey (ORCID 0009-0005-6869-308X) · published via DeskDeploy (deskdeploy.com), part of the Lattice… See the full description on the dataset page: https://huggingface.co/datasets/vincentcouey/standing-desk-claims-evidence-2026.botox-ad-hoc-claimsark-medical-claims-10202025claimsense-training-datatraining-dataset-for-proximal-claims-09-03-2024gym_claimstraining-dataset-for-30-juvederm-claims-07-29-2024claimsdataset
