datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
claim_stance
Dataset Card for Claim Stance Dataset
Dataset Summary
Claim Stance
This dataset contains 2,394 labeled Wikipedia claims for 55 topics. The dataset includes the stance (Pro/Con) of each claim towards the topic,
as well as fine-grained annotations, based on the semantic model of Stance Classification of Context-Dependent Claims (topic target,
topic sentiment towards its target, claim target, claim sentiment towards its target, and the relation between the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/claim_stance.Brevets-Francais-2025-Claimspatents_claims_1.5m_traim_testivermectin-cancer-claims
Ivermectin / Mebendazole / Fenbendazole — Cancer Case Claims
Private research dataset. Unverified patient claims extracted from public X posts by repurposed-drug advocates. NOT medical advice. NOT for clinical decision-making. Intended for signal-mining research.
Last updated: 2026-04-23
Contents
Two forms of the same data.
Parquet subsets (HF-native, loadable via datasets.load_dataset):
config
rows
description
patients
~540
one row per unique patient (dedup'd… See the full description on the dataset page: https://huggingface.co/datasets/erinkhoo/ivermectin-cancer-claims.claims_processing_results
Claims Processing — AM/UAM Evaluation Results
A courtesy release of the evaluation results produced for a master's thesis
on Polish-language claim verification / fact-checking. Shared in the interest of
research transparency and reproducibility.
Source code & full evaluation pipeline: https://github.com/Wilsonuep/claims_processing
License: MIT (see License below)
Language: Polish (pl)
All artifacts here were generated by running the agents defined in the repository
above… See the full description on the dataset page: https://huggingface.co/datasets/Wilsonuep/claims_processing_results.titer-expertise-claims
titer · attested expertise claims
99,984 expertise claims attested by publication record, for 20,000
researchers with an ORCID iD. Built to measure whether a people-search provider
can tell a real expert from a claimed one.
Sources: OpenAlex (disambiguated authors, works, topics), ORCID (the
identity spine), Crossref DOIs (the attestation chain). All open. A
researcher cannot self-assert a DOI into existence, which is why authorship is
treated as attested where a self-reported… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/titer-expertise-claims.AIPD_nlp_granted_claims
Dataset Card for "AIPD_nlp_granted_claims"
More Information needed
nlp-claims-processing
Nlp Claims Processing
Training dataset for NLP-based claims processing. Includes text classification and entity extraction.
Dataset Details
Records: 5,000
Features: 9
Organization: GCC Insurance ML Models Hub
Features
Column
Type
Description
claim_id
object
Feature for ML training
claim_text
object
Feature for ML training
category
object
Feature for ML training
sentiment
object
Feature for ML training
urgency_score
float64
Feature for ML… See the full description on the dataset page: https://huggingface.co/datasets/gcc-insurance-ml-models/nlp-claims-processing.emnlp-2020-2025-atomic-claims
EMNLP 2020–2025 Atomic Contribution Claims (ACC), with drift clusters
18,293 atomic contribution claims extracted from the abstracts of the full EMNLP main track 2020–2025, plus the canonical 80-cluster drift clustering and per-cluster drift statistics used in the Drift Inspector paper.
An atomic contribution claim (ACC) is a single self-contained sentence stating one concrete contribution of a paper: atomic (one contribution-bearing proposition), decontextualized (pronouns… See the full description on the dataset page: https://huggingface.co/datasets/Hamyrappy/emnlp-2020-2025-atomic-claims.fema-nfip-flood-insurance-claims
FEMA NFIP Redacted Claims v3 — free sample
This free 1,000-row sample spans all 49 loss years 1978–2026. The complete
ClarityStorm snapshot
contains 2,725,989 records through 2026-09-07, as of 2026-09-08, in CSV
and Parquet for $99 once. Future updates are not included. The underlying
FEMA source is free.
All 84 agency fields are preserved, plus total_paid_nominal, payment_status and
coordinate_status. Native camelCase names replace the old schema. The sample
selects evenly… See the full description on the dataset page: https://huggingface.co/datasets/claritystorm/fema-nfip-flood-insurance-claims.green_claims_annotatedFR-Patent-2025-Claimsclaim_stance
Dataset Card for Claim Stance Dataset
Dataset Summary
Claim Stance
This dataset contains 2,394 labeled Wikipedia claims for 55 topics. The dataset includes the stance (Pro/Con) of each claim towards the topic,
as well as fine-grained annotations, based on the semantic model of Stance Classification of Context-Dependent Claims (topic target,
topic sentiment towards its target, claim target, claim sentiment towards its target, and the relation… See the full description on the dataset page: https://huggingface.co/datasets/biityn/claim_stance.receipts-agent-claims
Receipts — Agent Claim Transcripts
Every transcript from the Receipts benchmark — one row per trial, graded by a pytest exit code rather than by another model.
424 runs on claude-haiku-4-5, plus 6 pilot runs on gemini-2.5-flash via aider. All trials are committed. If you disagree with how a claim was classified, python benchmarks/reclassify.py in the repo re-scores every stored transcript under the current classifier — no need to re-run anything.
What the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Hachiman94/receipts-agent-claims.green-claims-twitter-balancedafrica-synth-health-insurance-claims-all
African Health Insurance Claims | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-health-insurance-claims-all.presto-rumour-claims
PRESTO Corpus — Rumour Claims from 19th-Century British Newspapers
7,460 rumour claims automatically extracted from 19th-century British newspapers, plus a
200-row hand-adjudicated evaluation sample.
The source material is the British Library's Heritage Made Digital newspaper collection —
the same corpus published as biglam/hmd_newspapers.
PRESTO (Pattern-based Rumour Extraction with Semantic Tracking) applies dependency-pattern
matching to find and structure rumoured… See the full description on the dataset page: https://huggingface.co/datasets/biglam/presto-rumour-claims.insurance-motor-claims-decision-v1
Motor Insurance Claims Decision Support Dataset v1
Dataset Description
This is a synthetic dataset designed for training decision support models in motor insurance claims processing. It contains 800 records of motor insurance claims with associated decision labels (approve/review/reject).
Purpose: Human-in-the-loop claims decision support system training.
Schema
The dataset contains the following fields:
Field
Type
Description
claim_id
string
Unique… See the full description on the dataset page: https://huggingface.co/datasets/bdr-ai-org/insurance-motor-claims-decision-v1.French-Patent-2022-Claimsau-insurance-claims-quarterly
au-insurance-claims-quarterly
AU personal insurance claims register for 2026Q2 (general insurance). Each row is a policy-with-claim record including the policy premium.
Rows: 40
Columns: ['PolicyID', 'customer_id', 'product', 'premium_amount', 'claim_amt', 'claim_date', 'claim_status', 'region']
Brevets-Francais-2022-Claimsfever-claims-with-evidence-3classacl-anthology-atomic-claims
ACL Anthology Atomic Contribution Claims (ACC)
346,010 atomic contribution claims extracted from the abstracts of 80,144 ACL Anthology papers — 423 venues, 1963–2026.
An atomic contribution claim (ACC) is a single self-contained sentence stating one concrete contribution of a paper: atomic (exactly one contribution-bearing proposition), decontextualized (pronouns resolved, meta-language removed), and falsifiable (a verifiable assertion). Unlike raw abstracts or keywords, ACCs… See the full description on the dataset page: https://huggingface.co/datasets/Hamyrappy/acl-anthology-atomic-claims.za-egypt-insurance-claims-sample
za-egypt-insurance-claims-sample
A small sample of de-identified insurance claim records from South Africa (ZA) and Egypt (EG).
Dataset Summary
This dataset contains a sample of insurance claim records covering the South African and
Egyptian markets. It is intended as a lightweight example for exploring claims trends,
customer segmentation, and fraud-risk analysis. All records have been de-identified.
Total records: 80
Regions: South Africa (40), Egypt (40)… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/za-egypt-insurance-claims-sample.claims-risk-scoring-training
Claims Risk Scoring Training
Training dataset for claims risk scoring models. Includes severity predictions and payout estimates.
Dataset Details
Records: 12,000
Features: 17
Organization: GCC Insurance ML Models Hub
Features
Column
Type
Description
claim_id
object
Feature for ML training
claim_type
object
Feature for ML training
initial_estimate
float64
Feature for ML training
policy_coverage_limit
int64
Feature for ML training
deductible… See the full description on the dataset page: https://huggingface.co/datasets/gcc-insurance-ml-models/claims-risk-scoring-training.FR-Patent-2020-Claimsindian-insurance-claims-syntheticpatent-claims-hyaluronic-acid_enterprisepatent-claims-preservative_enterpriseBrevets-Francais-2020-Claims
