CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lucius-Morningstar /cms-desynpuf-insurance-claims CMS DE-SynPUF Insurance Claims (Rendered EOB Documents) Pipeline-ready insurance_claim evaluation corpus derived from the CMS 2008-2010 Data Entrepreneurs' Synthetic Public Use File (DE-SynPUF), Sample 1. One row = one Medicare claim rendered as a plain-text EOB-style document with ground-truth extraction fields aligned to the llm-mailroom InsuranceClaimExtraction schema. Provenance Source: CMS DE-SynPUF Sample 1 (fully synthetic; "very limited inferential… See the full description on the dataset page: https://huggingface.co/datasets/Lucius-Morningstar/cms-desynpuf-insurance-claims.texttoken-classificationn<1K2 likes171 downloads1mo agoHugging Face02referencesource /small-claims-court-dollar-limits-by-state Small Claims Court Dollar Limits by State Canonical, always-current version: https://referencesource.org/small-claims-court-dollar-limits-by-state/ Machine-readable: https://referencesource.org/small-claims-court-dollar-limits-by-state/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-26 Stale after: 2027-08-26 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 25 Every state sets its own maximum… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/small-claims-court-dollar-limits-by-state.textn<1K0 likes42 downloads29d agoHugging Face03Hachiman94 /receipts-agent-claims Receipts — Agent Claim Transcripts Every transcript from the Receipts benchmark — one row per trial, graded by a pytest exit code rather than by another model. 424 runs on claude-haiku-4-5, plus 6 pilot runs on gemini-2.5-flash via aider. All trials are committed. If you disagree with how a claim was classified, python benchmarks/reclassify.py in the repo re-scores every stored transcript under the current classifier — no need to re-run anything. What the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Hachiman94/receipts-agent-claims.tabulartext-generationn<1K0 likes42 downloads19d agoHugging Face04yjernite /ai-job-claimsClaims about AI's impact on jobs from experts quoted in the news articles collected in Florent Daudens' AI Job News dataset textn<1K0 likes40 downloads1y agoHugging Face05limmak /verified-step-claims-1k-preview Verified Step Claims 1K (Preview) Mathematically verified step-level claims with witness evidence for process supervision research. This is a 1K preview sampled from a larger in-progress corpus built from thousands of verified math solvers, covering Korean CSAT (수능), national mock exams, and advanced textbook problems across calculus, probability, sequences, and more. What Makes This Different PRM800K (OpenAI) This Dataset Label source Human annotators… See the full description on the dataset page: https://huggingface.co/datasets/limmak/verified-step-claims-1k-preview.text1K<n<10K0 likes39 downloads7mo agoHugging Face06ArtemkaT08 /medical-shorts-silver-claims-benchmark Medical Shorts Silver Claims Benchmark This dataset contains silver labels for evaluating medical claim detection and verification in short-form YouTube videos. The dataset does not include video files, audio files, thumbnails, or keyframes. It only includes YouTube video IDs, metadata, normalized medical claims, silver labels, and evidence source references. Dataset Size Videos: 236 Silver claims: 470 Visual-dependent claims: 22 Videos with medical claims: 205 Videos… See the full description on the dataset page: https://huggingface.co/datasets/ArtemkaT08/medical-shorts-silver-claims-benchmark.texttext-classificationn<1K0 likes35 downloads5mo agoHugging Face07agentlans /facebook-anli-claims Facebook ANLI Claims Dataset This dataset is a reformatted version of the facebook/anli dataset, reorganized by their original labels. The original dataset's splits have been combined and separated into three distinct dataset configurations based on their labels: entailment, neutral, and contradiction. For each configuration: Input: Premise text Output: A JSON list of hypotheses associated with the premise. text10K<n<100K0 likes30 downloads1y agoHugging Face08Teen-Different /grpo-oumi-synthetic-document-claims Dataset Card for GRPO Oumi ANLI Subset Dataset This dataset is a reformatted version of the oumi-ai/oumi-synthetic-document-claims dataset, specifically structured for use with the GRPO trainer. You can find more detailed information about the original dataset at the provided link. Link: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-document-claims Dataset Structure The dataset consists of a list of dictionaries, where each dictionary represents a… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-synthetic-document-claims.texttext-generation1K<n<10K0 likes26 downloads1y agoHugging Face09Loctran123 /vietnamese-fact-checking-claims Vietnamese Fact-Checking Claims Generated claim-verification data derived from the Vietnamese Evidence Corpus. Each article contains claims labeled as supported, refuted, or not having enough information, together with evidence and a short rationale. Statistics 12,238 source articles 73,454 generated claims 24,476 SUPPORTED claims 24,502 REFUTED claims 24,476 NOT_ENOUGH_INFO claims Main fields Article: id, date_iso, full_text, claims Claim: claim… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-fact-checking-claims.texttext-classification10K<n<100K0 likes20 downloads1mo agoHugging Face10Eugleo /echr-claimstext10K<n<100K0 likes15 downloads2y agoHugging Face11Teen-Different /grpo-oumi-synthetic-claims Dataset Card for GRPO Oumi ANLI Subset Dataset This dataset is a reformatted version of the TEEN-D/grpo-oumi-anli-subset dataset, specifically structured for use with the GRPO trainer. You can find more detailed information about the original dataset at the provided link. Link: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-claims Dataset Structure The dataset consists of a list of dictionaries, where each dictionary represents a single data instance… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-synthetic-claims.texttext-generation10K<n<100K0 likes15 downloads1y agoHugging Face12vincentcouey /standing-desk-claims-evidence-2026 DeskDeploy 2026 Standing Desk Health Claims vs Evidence Audit Brand-by-brand audit of 30 standing desk products mapping each marketed health claim to peer-reviewed evidence with Cochrane/PubMed evidence grades and verdicts. License: CC-BY 4.0 DOI: 10.5281/zenodo.20632751 Source study: https://deskdeploy.com/research/standing-desk-claims-evidence-2026/ Author: Vincent Wesley Couey (ORCID 0009-0005-6869-308X) · published via DeskDeploy (deskdeploy.com), part of the Lattice… See the full description on the dataset page: https://huggingface.co/datasets/vincentcouey/standing-desk-claims-evidence-2026.textn<1K0 likes12 downloads4mo agoHugging Face13Inabia-AI /botox-ad-hoc-claimstextn<1K0 likes11 downloads11mo agoHugging Face14Inabia-AI /ark-medical-claims-10202025text1K<n<10K0 likes4 downloads11mo agoHugging Face15pramodmisra /claimsense-training-datatext10K<n<100K0 likes4 downloads7mo agoHugging Face16Inabia-AI /training-dataset-for-proximal-claims-09-03-2024text1K<n<10K0 likes3 downloads2y agoHugging Face17DanRus21 /gym_claimstextn<1K0 likes3 downloads4mo agoHugging Face18Inabia-AI /training-dataset-for-30-juvederm-claims-07-29-2024textn<1K0 likes1 downloads2y agoHugging Face19Omkarrrkomarpant /claimsdatasettext1K<n<10K0 likes2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.