CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FactCheck-AI /FactCheck Dataset Card for FactCheck 📝 Dataset Summary FactCheck is an benchmark for evaluating LLMs on knowledge graph fact verification. It combines structured facts from YAGO, DBpedia, and FactBench with web-extracted evidence including questions, summaries, full text, and metadata. The dataset contains examples designed for sentence-level fact-checking and QA tasks. 📚 Supported Tasks Question Answering: Answer fact-checking questions derived from KG triples.… See the full description on the dataset page: https://huggingface.co/datasets/FactCheck-AI/FactCheck.imagequestion-answering1M<n<10M2 likes640 downloads8mo agoHugging Face02llm-semantic-router /fact-check-classification-dataset Fact-Check Classification Dataset 🎯 Purpose: Binary classification dataset for determining whether a prompt needs external fact-checking. Dataset Description This dataset is designed to train classifiers that can route LLM requests based on whether they require external fact verification. It's part of the vLLM Semantic Router project. Labels FACT_CHECK_NEEDED (1): Information-seeking questions requiring external verification Factual questions about dates… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/fact-check-classification-dataset.texttext-classification10K<n<100K2 likes213 downloads8mo agoHugging Face03datacommonsorg /datacommons_factcheckA dataset of fact checked claims by news media maintained by datacommons.orgtext-classification1K<n<10K5 likes151 downloads3y agoHugging Face04MultiMind-SemEval2025 /Augmented_MultiClaim_FactCheck_Retrieval Augmented MultiClaim FactCheck Retrieval Dataset 1. Dataset Summary This dataset is a collection of social media posts that have been augmented using a large language model (GPT-4o). The original dataset was sourced from the paper Multilingual Previously Fact-Checked Claim Retrieval by Matúš Pikuliak et al. (2023). You can access the original dataset from here. The dataset is used for improving the ability to comprehend content across multiple languages by integrating… See the full description on the dataset page: https://huggingface.co/datasets/MultiMind-SemEval2025/Augmented_MultiClaim_FactCheck_Retrieval.text10K<n<100K1 likes136 downloads1y agoHugging Face05xuandin /med-factcheck-benchmark-v20 likes132 downloads3mo agoHugging Face06tingchih /Multi_News_fact_checking_claims Dataset Card for "v2" More Information needed text1M<n<10M0 likes103 downloads3y agoHugging Face07SinclairSchneider /tweets_correctiv_and_factchecktabular1M<n<10M0 likes103 downloads4mo agoHugging Face08Cofacts /line-msg-fact-check-twgated Cofacts Archive for Reported Messages and Crowd-Sourced Fact-Check Replies The Cofacts dataset encompasses instant messages that have been reported by users of the Cofacts chatbot and the replies provided by the Cofacts crowd-sourced fact-checking community. Attribution to the Community This dataset is a result of contributions from both Cofacts LINE chatbot users and the community fact checkers. To appropriately attribute their efforts, please adhere to the… See the full description on the dataset page: https://huggingface.co/datasets/Cofacts/line-msg-fact-check-tw.tabulartext-classification10M<n<100M18 likes93 downloads4d agoHugging Face09Farsight-AI /10k-fact-check-finetune Dataset Card for "10k-fact-check-finetune" More Information needed text1K<n<10K0 likes91 downloads3y agoHugging Face10Ayush-242 /fact-checker-modeltext10K<n<100K0 likes86 downloads21d agoHugging Face11Lots-of-LoRAs /task966_ruletaker_fact_checking_based_on_given_context Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.texttext-generationn<1K0 likes85 downloads2y agoHugging Face12NaughtyConstrictor /fact-check-bureau Fact-Check Retrieval Dataset This dataset is designed to support the development and evaluation of fact-check retrieval pipelines. It is structured to work with FactCheckBureau, a tool for designing and evaluating fact-check retrieval pipelines. The dataset comprises a list of claims, fact-check articles, and precomputed embeddings for English and French fact-checks. Dataset Structure The dataset consists of the following files and directories: articles.csv: Contains… See the full description on the dataset page: https://huggingface.co/datasets/NaughtyConstrictor/fact-check-bureau.0 likes69 downloads2y agoHugging Face13xuandin /med-factcheck-benchmarktextn<1K0 likes62 downloads5mo agoHugging Face14Ayushnangia /moltbook-factcheck-conspiracy-grok Moltbook Factcheck Conspiracy — Grok Experiments Multi-agent social simulation data from Moltbook, a Reddit-like platform where AI agents autonomously post, comment, and vote. This dataset captures how Grok 4.1 Fast agents respond to conspiracy content seeded into their feed. Experiment Design Platform: Moltbook (Reddit-like social network for AI agents) Research Layer: CivicLens dose-response framework Duration: 1 hour per run Heartbeat: 60-second action cycle Date:… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-factcheck-conspiracy-grok.text-classificationn<1K1 likes60 downloads7mo agoHugging Face15ju-resplande /portuguese-fact-checking Portuguese Automated Fact-Checking Fake.BR COVID19.BR MuMiN-PT Info (fake/true) 🖥️ 💬 X Domain General Health "General" (Health) Year 2016–2018 2020 2020–2022 Approach [1] bottom-up bottom-up top-down Size 3580/3580 848/1139 1339/65 % URL 1.0%/0.7% 28.9%/56.9% 0.3%/0.0% Avg. # words 181.4/183.1 167.7/111.1 18.9/16.9 Corpora characteristics after cleaning. Top-down starts with fact-checked claims; bottom-up seeks for new misinformation in posts.… See the full description on the dataset page: https://huggingface.co/datasets/ju-resplande/portuguese-fact-checking.texttext-classification10K<n<100K0 likes56 downloads1y agoHugging Face16justinqbui /covid_fact_checked_google_apiThis dataset was gathered from the Google Fact Checker API, using an automatic web scraper. 10,000 facts were pulled, but for the sake of simplicity, only ones were the ratings were singular words "false" or "true", were kept, which filtered it down to ~3000 fact checks, with about 90% of the facts being false. annotations_creators: expert-generated language_creators: crowdsourced languages: en-US licenses: unknown multilinguality: monolingual pretty_name: polifact-covid-fact-checker… See the full description on the dataset page: https://huggingface.co/datasets/justinqbui/covid_fact_checked_google_api.text1K<n<10K1 likes52 downloads5y agoHugging Face17Ayushnangia /moltbook-factcheck-dose-response Moltbook Factcheck Dose-Response Experiment Multi-agent social simulation data from a dose-response experiment measuring how varying ratios of factual vs. conspiracy content affect AI agent behavior on a Reddit-like platform. Experiment Design Platform: Moltbook — a Reddit-like social network for AI agents Research layer: CivicLens — experiment infrastructure for controlled multi-agent studies Model: LLM-powered agents (10 agents per run + 2 system agents) Duration: ~1… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-factcheck-dose-response.text-classificationn<1K0 likes51 downloads7mo agoHugging Face18fake-news-UFG /FactChecksbrCollection of Portuguese Fact-Checking Benchmarks.text-classification10K<n<100K2 likes47 downloads3y agoHugging Face19Loctran123 /vietnamese-fact-checking-verifier-data Vietnamese fact-checking verifier data Leakage-aware document-level 80/10/10 split derived from Loctran123/vietnamese-fact-checking-claims at revision 63963b973af864ecc20313636b1786c31bbb4a41. Input is (evidence_text, claim) and labels are SUPPORTED, REFUTED, and NOT_ENOUGH_INFO. Exact duplicate claims are retained only once. texttext-classification10K<n<100K0 likes47 downloads1mo agoHugging Face20RotgarSett /chatgpt-clinic-caveats-fact-check When ChatGPT Adds Caveats to Clinic Recommendations This version 1.0 companion dataset fact-checks clinic-specific commercial and operational caveat families identified in a frozen corpus of 450 repeated ChatGPT answers from the parent study. Author: Evgeniy Yudin, Founder and Strategy Lead ORCID: https://orcid.org/0009-0007-8400-9561 Publisher: Rotgar Research Published: 2026-09-03 Version DOI: https://doi.org/10.5281/zenodo.22304469 Zenodo record:… See the full description on the dataset page: https://huggingface.co/datasets/RotgarSett/chatgpt-clinic-caveats-fact-check.text0 likes45 downloads20d agoHugging Face21Yoonseong /climatebert_factchecktext1K<n<10K4 likes39 downloads4y agoHugging Face22aiMy144 /viet-fact-checking Vietnamese Evidence Corpus for Fact-Checking & RAG (v1.0) This dataset is a clean, standardized, and unified Vietnamese Evidence Corpus (v1.0) built for research in Information Retrieval, Retrieval-Augmented Generation (RAG), and Fact-Checking / Claim Verification. Dataset Statistics Total Documents: 13,572 (frozen unique records, duplicates filtered out) Languages: ~70% Vietnamese (vi), ~30% English (en) Size: 115.33 MB Documents by Source… See the full description on the dataset page: https://huggingface.co/datasets/aiMy144/viet-fact-checking.texttext-retrieval10K<n<100K0 likes39 downloads2mo agoHugging Face23rimine /full_factchecktext10K<n<100K0 likes36 downloads1y agoHugging Face24sergiogpinto /factcheck-memes-x Fact-checking Memes - X Dataset This dataset contains 119 meme correction posts and their associated engagement metrics from a real-world deployment of fact-checking memes on X (formerly Twitter). The memes were specifically designed to counter misinformation by providing visually engaging explanations of fact-checking verdicts. Dataset Description Overview The "Fact-checking Memes - X" dataset documents a social media experiment conducted between October 25… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/factcheck-memes-x.imagen<1K0 likes32 downloads1y agoHugging Face25rickpereira /factguard_factchecking_datasetstext1K<n<10K0 likes32 downloads11mo agoHugging Face26Cyabra /ag_news_fact_check_with_llm Entity-Level Fact-Check Dataset Overview This dataset provides pairs of text snippets with controlled, entity-level factual perturbations, designed to evaluate large language models (LLMs) on their ability to detect, reason about, and correct factual errors at the entity level. Motivation Existing datasets (e.g., CNN/DailyMail, WikiBio, XSum) focus on broad factual consistency but do not provide explicit mappings between original facts and their incorrect… See the full description on the dataset page: https://huggingface.co/datasets/Cyabra/ag_news_fact_check_with_llm.texttext-classification1K<n<10K0 likes30 downloads1y agoHugging Face27drod75 /fact_checked_data_1text1M<n<10M0 likes30 downloads5mo agoHugging Face28jensjepsen /danish-sci-factcheck-v1tabular10K<n<100K0 likes28 downloads2mo agoHugging Face29ppaudel /twitter_factchecking_testtext1K<n<10K0 likes25 downloads2y agoHugging Face30Holmeister /TLT-FactChecktext1K<n<10K0 likes24 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.