datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-rule-checking
Arabic Rule Checking — قواعد ونصوص عربية بأحكام محسوبة
172,488 labelled (text, rule) pairs in Arabic. Each row asks one question: does this text
satisfy this rule? The answer is مطابق or مخالف.
بالعربية: مجموعة بيانات عربية للتحقق من مطابقة النصوص لقواعد مكتوبة بلغة طبيعية. كل صف
يحتوي على نص وقاعدة وحكم محسوب آليًا، وليس رأي نموذج.
split
pairs
texts
train
159,240
48,030
validation
13,248
2,002
Built from 50,062 generated Arabic texts across 12 document types… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rule-checking.viet-fact-checking
Vietnamese Evidence Corpus for Fact-Checking & RAG (v1.0)
This dataset is a clean, standardized, and unified Vietnamese Evidence Corpus (v1.0) built for research in Information Retrieval, Retrieval-Augmented Generation (RAG), and Fact-Checking / Claim Verification.
Dataset Statistics
Total Documents: 13,572 (frozen unique records, duplicates filtered out)
Languages: ~70% Vietnamese (vi), ~30% English (en)
Size: 115.33 MB
Documents by Source… See the full description on the dataset page: https://huggingface.co/datasets/aiMy144/viet-fact-checking.vietnamese-fact-checking-claims
Vietnamese Fact-Checking Claims
Generated claim-verification data derived from the Vietnamese Evidence Corpus.
Each article contains claims labeled as supported, refuted, or not having enough
information, together with evidence and a short rationale.
Statistics
12,238 source articles
73,454 generated claims
24,476 SUPPORTED claims
24,502 REFUTED claims
24,476 NOT_ENOUGH_INFO claims
Main fields
Article: id, date_iso, full_text, claims
Claim: claim… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-fact-checking-claims.resplit_multi_fact_checking_datasetfact-checking-embedding-datafact-checking-vietnamese-news
