datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multi_News_fact_checking_claims
Dataset Card for "v2"
More Information needed
arabic-rule-checking
Arabic Rule Checking — قواعد ونصوص عربية بأحكام محسوبة
172,488 labelled (text, rule) pairs in Arabic. Each row asks one question: does this text
satisfy this rule? The answer is مطابق or مخالف.
بالعربية: مجموعة بيانات عربية للتحقق من مطابقة النصوص لقواعد مكتوبة بلغة طبيعية. كل صف
يحتوي على نص وقاعدة وحكم محسوب آليًا، وليس رأي نموذج.
split
pairs
texts
train
159,240
48,030
validation
13,248
2,002
Built from 50,062 generated Arabic texts across 12 document types… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rule-checking.task966_ruletaker_fact_checking_based_on_given_context
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.checking
HumaniBench: A Human-Centric Visual QA Dataset
HumaniBench is a dataset for evaluating visual question answering models on tasks that involve human-centered attributes such as gender, age, and occupation.
Each data point includes:
ID: Unique identifier
Attribute: A social attribute (e.g., gender, race)
Question: A visual question related to the image
Answer: The ground-truth answer
image: Embedded image in base64 or file format for visual preview
Example Entry
{… See the full description on the dataset page: https://huggingface.co/datasets/shainaraza/checking.vietnamese-fact-checking-verifier-data
Vietnamese fact-checking verifier data
Leakage-aware document-level 80/10/10 split derived from
Loctran123/vietnamese-fact-checking-claims at revision 63963b973af864ecc20313636b1786c31bbb4a41.
Input is (evidence_text, claim) and labels are SUPPORTED, REFUTED, and
NOT_ENOUGH_INFO. Exact duplicate claims are retained only once.
portuguese-fact-checking
Portuguese Automated Fact-Checking
Fake.BR
COVID19.BR
MuMiN-PT
Info (fake/true)
🖥️
💬
X
Domain
General
Health
"General" (Health)
Year
2016–2018
2020
2020–2022
Approach [1]
bottom-up
bottom-up
top-down
Size
3580/3580
848/1139
1339/65
% URL
1.0%/0.7%
28.9%/56.9%
0.3%/0.0%
Avg. # words
181.4/183.1
167.7/111.1
18.9/16.9
Corpora characteristics after cleaning. Top-down starts with fact-checked claims; bottom-up seeks for new misinformation in posts.… See the full description on the dataset page: https://huggingface.co/datasets/ju-resplande/portuguese-fact-checking.viet-fact-checking
Vietnamese Evidence Corpus for Fact-Checking & RAG (v1.0)
This dataset is a clean, standardized, and unified Vietnamese Evidence Corpus (v1.0) built for research in Information Retrieval, Retrieval-Augmented Generation (RAG), and Fact-Checking / Claim Verification.
Dataset Statistics
Total Documents: 13,572 (frozen unique records, duplicates filtered out)
Languages: ~70% Vietnamese (vi), ~30% English (en)
Size: 115.33 MB
Documents by Source… See the full description on the dataset page: https://huggingface.co/datasets/aiMy144/viet-fact-checking.Scripts_for_checking_Train.pyThis includes python script (will run in CMD console in Windows 10 local computer) to EXTRACT the words from a Train_Text.txt file to examine and extract the whole-words incuded in the text.
A second python script sorts the Extracted words into 1-letter_words.txt, 2-letters_words.txt, 3-letters_words.txt, and 4andmore-letter_words.txt, plus unicode(Chinese)_words.txt, numeralized_words.txt any Any_other_words.txt
P.S. I noticed in November 2024 that: The python script considers an underbar _… See the full description on the dataset page: https://huggingface.co/datasets/MartialTerran/Scripts_for_checking_Train.py.wan_checking_watchThis dataset contains videos generated using Wan 2.1 T2V 14B.
vietnamese-fact-checking-verifier-data-v3-1
Vietnamese fact-checking verifier data
Leakage-aware document-level 80/10/10 split derived from
aiMy144/vietnamese-fact-checking-claims-v3-1 at revision e20c1eddcbb4a5de862dbc6965eabee43c9766da.
Input is (evidence_text, claim) and labels are SUPPORTED, REFUTED, and
NOT_ENOUGH_INFO. Exact duplicate claims are retained only once.
fact-checkingvietnamese-fact-checking-claims
Vietnamese Fact-Checking Claims
Generated claim-verification data derived from the Vietnamese Evidence Corpus.
Each article contains claims labeled as supported, refuted, or not having enough
information, together with evidence and a short rationale.
Statistics
12,238 source articles
73,454 generated claims
24,476 SUPPORTED claims
24,502 REFUTED claims
24,476 NOT_ENOUGH_INFO claims
Main fields
Article: id, date_iso, full_text, claims
Claim: claim… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-fact-checking-claims.condition-checking-dataset
Condition Checking Dataset
This dataset contains condition checking conversations for robotics applications, with embedded base64 images from multiple camera viewpoints.
Dataset Structure
Data Fields
id: Unique identifier for each sample
images: Dictionary containing base64-encoded images from multiple camera viewpoints
conversations: List of conversation turns (human question + assistant answer)
Camera Viewpoints
The dataset includes images from 5… See the full description on the dataset page: https://huggingface.co/datasets/jeffshen4011/condition-checking-dataset.vietnamese-fact-checking-claims-v3-1
Vietnamese Fact-Checking Claims v3.1
Generated three-label claims associated with the Vietnamese evidence corpus.
Statistics
Total claims: 73,454
SUPPORTED: 24,477
REFUTED: 24,501
NOT_ENOUGH_INFO: 24,476
Normalized unique claims: 72,863
Missing or unindexable evidence document IDs: 0
Empty evidence quotes: 0
Version 3.1 repairs 40 evidence entries across 35 claims. Thirty-three entries
were missing article_id; seven contained typos or corrupted IDs. Each repair… See the full description on the dataset page: https://huggingface.co/datasets/aiMy144/vietnamese-fact-checking-claims-v3-1.resplit_multi_fact_checking_datasetchecking-test
Dataset Card for "checking-test"
More Information needed
rlvr_task966_ruletaker_fact_checking_based_on_given_contextDisaster-Type_Classification_Dataset_for_Automated_Fact-Checking
DTCD-AFC: Disaster-Type Classification Dataset for Automated Fact-Checking
Overview
The DTCD-AFC is a dataset designed for disaster-type classification evaluation for automated fact-checking.
It consists of multimodal social media posts collected based on past natural disasters, each labeled with the disaster type to which its content relates.
The social media posts are sourced from CrisisMMD.
Files
disaster_type_classification_dataset_for_afc.csv: The CSV… See the full description on the dataset page: https://huggingface.co/datasets/o-yas/Disaster-Type_Classification_Dataset_for_Automated_Fact-Checking.flan_combined_task966_ruletaker_fact_checking_based_on_given_contextultra-feedback_checkingqwen3_0.6b-rlvr_task966_ruletaker_fact_checking_based_on_given_contextfact-checking-embedding-datachecking_medicalscheckingfact-checking-vietnamese-newscompliance_checkingcheckingCheckingcuda-stack-casssubset-checkingdZhabDtALkOxnJYluQUrCBvRrGd2_checking_bug
