datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
claim_stance
Dataset Card for Claim Stance Dataset
Dataset Summary
Claim Stance
This dataset contains 2,394 labeled Wikipedia claims for 55 topics. The dataset includes the stance (Pro/Con) of each claim towards the topic,
as well as fine-grained annotations, based on the semantic model of Stance Classification of Context-Dependent Claims (topic target,
topic sentiment towards its target, claim target, claim sentiment towards its target, and the relation between the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/claim_stance.patents_claims_1.5m_traim_testinsurance-claims-extractionThis dataset can be used for benchmarking LLM Structured Outputs via the code here:
https://github.com/cleanlab/structured-output-benchmark/
fema-nfip-flood-insurance-claims
FEMA NFIP Redacted Claims v3 — free sample
This free 1,000-row sample spans all 49 loss years 1978–2026. The complete
ClarityStorm snapshot
contains 2,725,989 records through 2026-09-07, as of 2026-09-08, in CSV
and Parquet for $99 once. Future updates are not included. The underlying
FEMA source is free.
All 84 agency fields are preserved, plus total_paid_nominal, payment_status and
coordinate_status. Native camelCase names replace the old schema. The sample
selects evenly… See the full description on the dataset page: https://huggingface.co/datasets/claritystorm/fema-nfip-flood-insurance-claims.green_claims_annotatedclaim_stance
Dataset Card for Claim Stance Dataset
Dataset Summary
Claim Stance
This dataset contains 2,394 labeled Wikipedia claims for 55 topics. The dataset includes the stance (Pro/Con) of each claim towards the topic,
as well as fine-grained annotations, based on the semantic model of Stance Classification of Context-Dependent Claims (topic target,
topic sentiment towards its target, claim target, claim sentiment towards its target, and the relation… See the full description on the dataset page: https://huggingface.co/datasets/biityn/claim_stance.HiCaD_claims
References
Hierarchical Catalogue Generation for Literature Review: A Benchmark
Authors: Kun Zhu, Xiaocheng Feng, Xiachong Feng, Yingsheng Wu, Bing Qin
Abstract: This paper presents a benchmark for hierarchical catalogue generation to aid literature reviews, addressing the challenge of organizing information from numerous references into a coherent structure.
Findings: EMNLP 2023
Link: arXiv:2304.03512
CHIME: LLM-Assisted Hierarchical Organization of… See the full description on the dataset page: https://huggingface.co/datasets/technicolor/HiCaD_claims.HiCaD_claims_with_prompt
References
Hierarchical Catalogue Generation for Literature Review: A Benchmark
Authors: Kun Zhu, Xiaocheng Feng, Xiachong Feng, Yingsheng Wu, Bing Qin
Abstract: This paper presents a benchmark for hierarchical catalogue generation to aid literature reviews, addressing the challenge of organizing information from numerous references into a coherent structure.
Findings: EMNLP 2023
Link: arXiv:2304.03512
CHIME: LLM-Assisted Hierarchical Organization of… See the full description on the dataset page: https://huggingface.co/datasets/technicolor/HiCaD_claims_with_prompt.za-egypt-insurance-claims-sample
za-egypt-insurance-claims-sample
A small sample of de-identified insurance claim records from South Africa (ZA) and Egypt (EG).
Dataset Summary
This dataset contains a sample of insurance claim records covering the South African and
Egyptian markets. It is intended as a lightweight example for exploring claims trends,
customer segmentation, and fraud-risk analysis. All records have been de-identified.
Total records: 80
Regions: South Africa (40), Egypt (40)… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/za-egypt-insurance-claims-sample.claim_stance
Dataset Card for Claim Stance Dataset
Dataset Summary
Claim Stance
This dataset contains 2,394 labeled Wikipedia claims for 55 topics. The dataset includes the stance (Pro/Con) of each claim towards the topic,
as well as fine-grained annotations, based on the semantic model of Stance Classification of Context-Dependent Claims (topic target,
topic sentiment towards its target, claim target, claim sentiment towards its target, and the relation between the… See the full description on the dataset page: https://huggingface.co/datasets/CHAMP12154/claim_stance.contrarian-claimsspecialty-clinic-synthetic-claims
Specialty Clinic Synthetic Claims Sample
This sample contains 1000 generated-from-scratch synthetic healthcare claim rows.
Use case: Model specialty clinic procedure authorization, documentation, and payer coverage friction across multi-specialty outpatient settings.
Synthetic-Only Boundary
No PHI.
No customer data.
No patient-level source records.
Generated from public healthcare references, explicit modeling assumptions, and deterministic
synthesis code.… See the full description on the dataset page: https://huggingface.co/datasets/Upstream-Intelligence/specialty-clinic-synthetic-claims.green-claims-twitterHiCaD_claims_with_prompt_jsonclaims_1000description_to_claims_splitclaim-span-datasetsingapore-misinformation-claims-with-evidencedescription_to_claimsval_description_to_claimspatents_claims_1.5m_traim_testclaims
