datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FactCheck
Dataset Card for FactCheck
📝 Dataset Summary
FactCheck is an benchmark for evaluating LLMs on knowledge graph fact verification. It combines structured facts from YAGO, DBpedia, and FactBench with web-extracted evidence including questions, summaries, full text, and metadata. The dataset contains examples designed for sentence-level fact-checking and QA tasks.
📚 Supported Tasks
Question Answering: Answer fact-checking questions derived from KG triples.… See the full description on the dataset page: https://huggingface.co/datasets/FactCheck-AI/FactCheck.fact-check-classification-dataset
Fact-Check Classification Dataset
🎯 Purpose: Binary classification dataset for determining whether a prompt needs external fact-checking.
Dataset Description
This dataset is designed to train classifiers that can route LLM requests based on whether they require external fact verification. It's part of the vLLM Semantic Router project.
Labels
FACT_CHECK_NEEDED (1): Information-seeking questions requiring external verification
Factual questions about dates… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/fact-check-classification-dataset.Augmented_MultiClaim_FactCheck_Retrieval
Augmented MultiClaim FactCheck Retrieval Dataset
1. Dataset Summary
This dataset is a collection of social media posts that have been augmented using a large language model (GPT-4o). The original dataset was sourced from the paper Multilingual Previously Fact-Checked Claim Retrieval by Matúš Pikuliak et al. (2023). You can access the original dataset from here. The dataset is used for improving the ability to comprehend content across multiple languages by integrating… See the full description on the dataset page: https://huggingface.co/datasets/MultiMind-SemEval2025/Augmented_MultiClaim_FactCheck_Retrieval.Multi_News_fact_checking_claims
Dataset Card for "v2"
More Information needed
tweets_correctiv_and_factcheckline-msg-fact-check-tw
Cofacts Archive for Reported Messages and Crowd-Sourced Fact-Check Replies
The Cofacts dataset encompasses instant messages that have been reported by users of the Cofacts chatbot and the replies provided by the Cofacts crowd-sourced fact-checking community.
Attribution to the Community
This dataset is a result of contributions from both Cofacts LINE chatbot users and the community fact checkers.
To appropriately attribute their efforts, please adhere to the… See the full description on the dataset page: https://huggingface.co/datasets/Cofacts/line-msg-fact-check-tw.10k-fact-check-finetune
Dataset Card for "10k-fact-check-finetune"
More Information needed
fact-checker-modeltask966_ruletaker_fact_checking_based_on_given_context
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.med-factcheck-benchmarkportuguese-fact-checking
Portuguese Automated Fact-Checking
Fake.BR
COVID19.BR
MuMiN-PT
Info (fake/true)
🖥️
💬
X
Domain
General
Health
"General" (Health)
Year
2016–2018
2020
2020–2022
Approach [1]
bottom-up
bottom-up
top-down
Size
3580/3580
848/1139
1339/65
% URL
1.0%/0.7%
28.9%/56.9%
0.3%/0.0%
Avg. # words
181.4/183.1
167.7/111.1
18.9/16.9
Corpora characteristics after cleaning. Top-down starts with fact-checked claims; bottom-up seeks for new misinformation in posts.… See the full description on the dataset page: https://huggingface.co/datasets/ju-resplande/portuguese-fact-checking.covid_fact_checked_google_apiThis dataset was gathered from the Google Fact Checker API, using an automatic web scraper. 10,000 facts were pulled, but for the sake of simplicity, only ones were the ratings were singular words "false" or "true", were kept, which filtered it down to ~3000 fact checks, with about 90% of the facts being false.
annotations_creators:
expert-generated
language_creators:
crowdsourced
languages:
en-US
licenses:
unknown
multilinguality:
monolingual
pretty_name: polifact-covid-fact-checker… See the full description on the dataset page: https://huggingface.co/datasets/justinqbui/covid_fact_checked_google_api.vietnamese-fact-checking-verifier-data
Vietnamese fact-checking verifier data
Leakage-aware document-level 80/10/10 split derived from
Loctran123/vietnamese-fact-checking-claims at revision 63963b973af864ecc20313636b1786c31bbb4a41.
Input is (evidence_text, claim) and labels are SUPPORTED, REFUTED, and
NOT_ENOUGH_INFO. Exact duplicate claims are retained only once.
chatgpt-clinic-caveats-fact-check
When ChatGPT Adds Caveats to Clinic Recommendations
This version 1.0 companion dataset fact-checks clinic-specific commercial and operational caveat families identified in a frozen corpus of 450 repeated ChatGPT answers from the parent study.
Author: Evgeniy Yudin, Founder and Strategy Lead
ORCID: https://orcid.org/0009-0007-8400-9561
Publisher: Rotgar Research
Published: 2026-09-03
Version DOI: https://doi.org/10.5281/zenodo.22304469
Zenodo record:… See the full description on the dataset page: https://huggingface.co/datasets/RotgarSett/chatgpt-clinic-caveats-fact-check.climatebert_factcheckviet-fact-checking
Vietnamese Evidence Corpus for Fact-Checking & RAG (v1.0)
This dataset is a clean, standardized, and unified Vietnamese Evidence Corpus (v1.0) built for research in Information Retrieval, Retrieval-Augmented Generation (RAG), and Fact-Checking / Claim Verification.
Dataset Statistics
Total Documents: 13,572 (frozen unique records, duplicates filtered out)
Languages: ~70% Vietnamese (vi), ~30% English (en)
Size: 115.33 MB
Documents by Source… See the full description on the dataset page: https://huggingface.co/datasets/aiMy144/viet-fact-checking.full_factcheckfactcheck-memes-x
Fact-checking Memes - X Dataset
This dataset contains 119 meme correction posts and their associated engagement metrics from a real-world deployment of fact-checking memes on X (formerly Twitter). The memes were specifically designed to counter misinformation by providing visually engaging explanations of fact-checking verdicts.
Dataset Description
Overview
The "Fact-checking Memes - X" dataset documents a social media experiment conducted between October 25… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/factcheck-memes-x.factguard_factchecking_datasetsag_news_fact_check_with_llm
Entity-Level Fact-Check Dataset
Overview
This dataset provides pairs of text snippets with controlled, entity-level factual perturbations, designed to evaluate large language models (LLMs) on their ability to detect, reason about, and correct factual errors at the entity level.
Motivation
Existing datasets (e.g., CNN/DailyMail, WikiBio, XSum) focus on broad factual consistency but do not provide explicit mappings between original facts and their incorrect… See the full description on the dataset page: https://huggingface.co/datasets/Cyabra/ag_news_fact_check_with_llm.fact_checked_data_1danish-sci-factcheck-v1twitter_factchecking_testTLT-FactCheckFactCheck-multifactcheck_datasetFactCheck-singleThis dataset is a combination of originally English Climate-Fever, SciFact, COVID-Fact, and HealthVer datasets' translation to Turkish. Translation is done automatically,
so there can be inaccurate translations. Each test instance is paired with 10 different instructions for multi-prompt evaluation.
Original Datasets
Diggelmann, Thomas; Boyd-Graber, Jordan; Bulian, Jannis; Ciaramita, Massimiliano; Leippold, Markus (2020). CLIMATE-FEVER: A Dataset for Verification of Real-World… See the full description on the dataset page: https://huggingface.co/datasets/Holmeister/FactCheck-single.factcheck_srvietnamese-fact-checking-claims
Vietnamese Fact-Checking Claims
Generated claim-verification data derived from the Vietnamese Evidence Corpus.
Each article contains claims labeled as supported, refuted, or not having enough
information, together with evidence and a short rationale.
Statistics
12,238 source articles
73,454 generated claims
24,476 SUPPORTED claims
24,502 REFUTED claims
24,476 NOT_ENOUGH_INFO claims
Main fields
Article: id, date_iso, full_text, claims
Claim: claim… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-fact-checking-claims.covid_fact_checked_polifactThis dataset was gathered by using an automated web scraper that scraped polifact covid fact checker. This dataset contains three columns, the text, the rating given by polifact (half-true, full-flop, pants-fire, barely-true true, mostly-true, and false), and the adjusted rating.
The adjusted rating was created by mapping the raw rating given by polifact
true -> true
mostly-true -> true
half-true -> misleading
barely-true -> misleading
false -> false
pants-fire -> false
full-flop -> false… See the full description on the dataset page: https://huggingface.co/datasets/justinqbui/covid_fact_checked_polifact.
