CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FactCheck-AI /FactCheck Dataset Card for FactCheck 📝 Dataset Summary FactCheck is an benchmark for evaluating LLMs on knowledge graph fact verification. It combines structured facts from YAGO, DBpedia, and FactBench with web-extracted evidence including questions, summaries, full text, and metadata. The dataset contains examples designed for sentence-level fact-checking and QA tasks. 📚 Supported Tasks Question Answering: Answer fact-checking questions derived from KG triples.… See the full description on the dataset page: https://huggingface.co/datasets/FactCheck-AI/FactCheck.imagequestion-answering1M<n<10M2 likes640 downloads8mo agoHugging Face02llm-semantic-router /fact-check-classification-dataset Fact-Check Classification Dataset 🎯 Purpose: Binary classification dataset for determining whether a prompt needs external fact-checking. Dataset Description This dataset is designed to train classifiers that can route LLM requests based on whether they require external fact verification. It's part of the vLLM Semantic Router project. Labels FACT_CHECK_NEEDED (1): Information-seeking questions requiring external verification Factual questions about dates… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/fact-check-classification-dataset.texttext-classification10K<n<100K2 likes213 downloads8mo agoHugging Face03MultiMind-SemEval2025 /Augmented_MultiClaim_FactCheck_Retrieval Augmented MultiClaim FactCheck Retrieval Dataset 1. Dataset Summary This dataset is a collection of social media posts that have been augmented using a large language model (GPT-4o). The original dataset was sourced from the paper Multilingual Previously Fact-Checked Claim Retrieval by Matúš Pikuliak et al. (2023). You can access the original dataset from here. The dataset is used for improving the ability to comprehend content across multiple languages by integrating… See the full description on the dataset page: https://huggingface.co/datasets/MultiMind-SemEval2025/Augmented_MultiClaim_FactCheck_Retrieval.text10K<n<100K1 likes136 downloads1y agoHugging Face04tingchih /Multi_News_fact_checking_claims Dataset Card for "v2" More Information needed text1M<n<10M0 likes103 downloads3y agoHugging Face05SinclairSchneider /tweets_correctiv_and_factchecktabular1M<n<10M0 likes103 downloads4mo agoHugging Face06Cofacts /line-msg-fact-check-twgated Cofacts Archive for Reported Messages and Crowd-Sourced Fact-Check Replies The Cofacts dataset encompasses instant messages that have been reported by users of the Cofacts chatbot and the replies provided by the Cofacts crowd-sourced fact-checking community. Attribution to the Community This dataset is a result of contributions from both Cofacts LINE chatbot users and the community fact checkers. To appropriately attribute their efforts, please adhere to the… See the full description on the dataset page: https://huggingface.co/datasets/Cofacts/line-msg-fact-check-tw.tabulartext-classification10M<n<100M18 likes93 downloads4d agoHugging Face07Farsight-AI /10k-fact-check-finetune Dataset Card for "10k-fact-check-finetune" More Information needed text1K<n<10K0 likes91 downloads3y agoHugging Face08Ayush-242 /fact-checker-modeltext10K<n<100K0 likes86 downloads21d agoHugging Face09Lots-of-LoRAs /task966_ruletaker_fact_checking_based_on_given_context Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.texttext-generationn<1K0 likes85 downloads2y agoHugging Face10xuandin /med-factcheck-benchmarktextn<1K0 likes62 downloads5mo agoHugging Face11ju-resplande /portuguese-fact-checking Portuguese Automated Fact-Checking Fake.BR COVID19.BR MuMiN-PT Info (fake/true) 🖥️ 💬 X Domain General Health "General" (Health) Year 2016–2018 2020 2020–2022 Approach [1] bottom-up bottom-up top-down Size 3580/3580 848/1139 1339/65 % URL 1.0%/0.7% 28.9%/56.9% 0.3%/0.0% Avg. # words 181.4/183.1 167.7/111.1 18.9/16.9 Corpora characteristics after cleaning. Top-down starts with fact-checked claims; bottom-up seeks for new misinformation in posts.… See the full description on the dataset page: https://huggingface.co/datasets/ju-resplande/portuguese-fact-checking.texttext-classification10K<n<100K0 likes56 downloads1y agoHugging Face12justinqbui /covid_fact_checked_google_apiThis dataset was gathered from the Google Fact Checker API, using an automatic web scraper. 10,000 facts were pulled, but for the sake of simplicity, only ones were the ratings were singular words "false" or "true", were kept, which filtered it down to ~3000 fact checks, with about 90% of the facts being false. annotations_creators: expert-generated language_creators: crowdsourced languages: en-US licenses: unknown multilinguality: monolingual pretty_name: polifact-covid-fact-checker… See the full description on the dataset page: https://huggingface.co/datasets/justinqbui/covid_fact_checked_google_api.text1K<n<10K1 likes52 downloads5y agoHugging Face13Loctran123 /vietnamese-fact-checking-verifier-data Vietnamese fact-checking verifier data Leakage-aware document-level 80/10/10 split derived from Loctran123/vietnamese-fact-checking-claims at revision 63963b973af864ecc20313636b1786c31bbb4a41. Input is (evidence_text, claim) and labels are SUPPORTED, REFUTED, and NOT_ENOUGH_INFO. Exact duplicate claims are retained only once. texttext-classification10K<n<100K0 likes47 downloads1mo agoHugging Face14RotgarSett /chatgpt-clinic-caveats-fact-check When ChatGPT Adds Caveats to Clinic Recommendations This version 1.0 companion dataset fact-checks clinic-specific commercial and operational caveat families identified in a frozen corpus of 450 repeated ChatGPT answers from the parent study. Author: Evgeniy Yudin, Founder and Strategy Lead ORCID: https://orcid.org/0009-0007-8400-9561 Publisher: Rotgar Research Published: 2026-09-03 Version DOI: https://doi.org/10.5281/zenodo.22304469 Zenodo record:… See the full description on the dataset page: https://huggingface.co/datasets/RotgarSett/chatgpt-clinic-caveats-fact-check.text0 likes45 downloads21d agoHugging Face15Yoonseong /climatebert_factchecktext1K<n<10K4 likes39 downloads4y agoHugging Face16aiMy144 /viet-fact-checking Vietnamese Evidence Corpus for Fact-Checking & RAG (v1.0) This dataset is a clean, standardized, and unified Vietnamese Evidence Corpus (v1.0) built for research in Information Retrieval, Retrieval-Augmented Generation (RAG), and Fact-Checking / Claim Verification. Dataset Statistics Total Documents: 13,572 (frozen unique records, duplicates filtered out) Languages: ~70% Vietnamese (vi), ~30% English (en) Size: 115.33 MB Documents by Source… See the full description on the dataset page: https://huggingface.co/datasets/aiMy144/viet-fact-checking.texttext-retrieval10K<n<100K0 likes39 downloads2mo agoHugging Face17rimine /full_factchecktext10K<n<100K0 likes36 downloads1y agoHugging Face18sergiogpinto /factcheck-memes-x Fact-checking Memes - X Dataset This dataset contains 119 meme correction posts and their associated engagement metrics from a real-world deployment of fact-checking memes on X (formerly Twitter). The memes were specifically designed to counter misinformation by providing visually engaging explanations of fact-checking verdicts. Dataset Description Overview The "Fact-checking Memes - X" dataset documents a social media experiment conducted between October 25… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/factcheck-memes-x.imagen<1K0 likes32 downloads1y agoHugging Face19rickpereira /factguard_factchecking_datasetstext1K<n<10K0 likes32 downloads11mo agoHugging Face20Cyabra /ag_news_fact_check_with_llm Entity-Level Fact-Check Dataset Overview This dataset provides pairs of text snippets with controlled, entity-level factual perturbations, designed to evaluate large language models (LLMs) on their ability to detect, reason about, and correct factual errors at the entity level. Motivation Existing datasets (e.g., CNN/DailyMail, WikiBio, XSum) focus on broad factual consistency but do not provide explicit mappings between original facts and their incorrect… See the full description on the dataset page: https://huggingface.co/datasets/Cyabra/ag_news_fact_check_with_llm.texttext-classification1K<n<10K0 likes30 downloads1y agoHugging Face21drod75 /fact_checked_data_1text1M<n<10M0 likes30 downloads5mo agoHugging Face22jensjepsen /danish-sci-factcheck-v1tabular10K<n<100K0 likes28 downloads2mo agoHugging Face23ppaudel /twitter_factchecking_testtext1K<n<10K0 likes25 downloads2y agoHugging Face24Holmeister /TLT-FactChecktext1K<n<10K0 likes24 downloads2y agoHugging Face25Holmeister /FactCheck-multitext10K<n<100K0 likes24 downloads2y agoHugging Face26rimine /factcheck_datasettext10K<n<100K0 likes23 downloads1y agoHugging Face27Holmeister /FactCheck-singleThis dataset is a combination of originally English Climate-Fever, SciFact, COVID-Fact, and HealthVer datasets' translation to Turkish. Translation is done automatically, so there can be inaccurate translations. Each test instance is paired with 10 different instructions for multi-prompt evaluation. Original Datasets Diggelmann, Thomas; Boyd-Graber, Jordan; Bulian, Jannis; Ciaramita, Massimiliano; Leippold, Markus (2020). CLIMATE-FEVER: A Dataset for Verification of Real-World… See the full description on the dataset page: https://huggingface.co/datasets/Holmeister/FactCheck-single.text10K<n<100K0 likes22 downloads2y agoHugging Face28rimine /factcheck_srtext10K<n<100K0 likes20 downloads1y agoHugging Face29Loctran123 /vietnamese-fact-checking-claims Vietnamese Fact-Checking Claims Generated claim-verification data derived from the Vietnamese Evidence Corpus. Each article contains claims labeled as supported, refuted, or not having enough information, together with evidence and a short rationale. Statistics 12,238 source articles 73,454 generated claims 24,476 SUPPORTED claims 24,502 REFUTED claims 24,476 NOT_ENOUGH_INFO claims Main fields Article: id, date_iso, full_text, claims Claim: claim… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-fact-checking-claims.texttext-classification10K<n<100K0 likes20 downloads1mo agoHugging Face30justinqbui /covid_fact_checked_polifactThis dataset was gathered by using an automated web scraper that scraped polifact covid fact checker. This dataset contains three columns, the text, the rating given by polifact (half-true, full-flop, pants-fire, barely-true true, mostly-true, and false), and the adjusted rating. The adjusted rating was created by mapping the raw rating given by polifact true -> true mostly-true -> true half-true -> misleading barely-true -> misleading false -> false pants-fire -> false full-flop -> false… See the full description on the dataset page: https://huggingface.co/datasets/justinqbui/covid_fact_checked_polifact.text1K<n<10K3 likes17 downloads5y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.