datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
med-factcheck-benchmarkviet-fact-checking
Vietnamese Evidence Corpus for Fact-Checking & RAG (v1.0)
This dataset is a clean, standardized, and unified Vietnamese Evidence Corpus (v1.0) built for research in Information Retrieval, Retrieval-Augmented Generation (RAG), and Fact-Checking / Claim Verification.
Dataset Statistics
Total Documents: 13,572 (frozen unique records, duplicates filtered out)
Languages: ~70% Vietnamese (vi), ~30% English (en)
Size: 115.33 MB
Documents by Source… See the full description on the dataset page: https://huggingface.co/datasets/aiMy144/viet-fact-checking.vietnamese-fact-checking-claims
Vietnamese Fact-Checking Claims
Generated claim-verification data derived from the Vietnamese Evidence Corpus.
Each article contains claims labeled as supported, refuted, or not having enough
information, together with evidence and a short rationale.
Statistics
12,238 source articles
73,454 generated claims
24,476 SUPPORTED claims
24,502 REFUTED claims
24,476 NOT_ENOUGH_INFO claims
Main fields
Article: id, date_iso, full_text, claims
Claim: claim… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-fact-checking-claims.resplit_multi_fact_checking_datasetscience-factcheck-indic
Health-Science Fact-Check (Hindi/Punjabi)
Native Hindi and Punjabi text from
ai4bharat/IndicCorpV2,
adapted with AutoScientist into substantive domain responses written by
a health-science fact-checker judging claims and showing the reasoning.
Rows
667
Unique source texts
667
Absolute quality score
9.0/10 (grade A)
Source score before adaptation
9.0/10 (grade A)
Percentile
33.0
Relative change
+0.0%
Median response length
520 chars… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/science-factcheck-indic.fact-checking-embedding-dataveritas-factcheck-datafact-checking-vietnamese-news
