CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /FACTS-grounding-public FACTS Grounding 1.0 Public Examples 860 public FACTS Grounding examples from Google DeepMind and Google Research FACTS Grounding is a benchmark from Google DeepMind and Google Research designed to measure the performance of AI Models on factuality and grounding. ▶ FACTS Grounding Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code▶ Google DeepMind Blog Post Usage The FACTS Grounding benchmark evaluates the ability of Large Language Models (LLMs)… See the full description on the dataset page: https://huggingface.co/datasets/google/FACTS-grounding-public.textquestion-answeringn<1K47 likes1.2k downloads2y agoHugging Face02dynamicfeed /live-facts-snapshot Live Facts Snapshot A daily snapshot of verifiable, post-training-cutoff world-state facts — the kind of ground truth language models cannot know from training data — exported through Dynamic Feed, a live, verifiable data API whose every response is Ed25519-signed. One file per day (data/YYYY-MM-DD.jsonl), one fact per line, and every row carries its own source, source_url and measured_at. Facts covered per day: tool facts upstream source licence software_version… See the full description on the dataset page: https://huggingface.co/datasets/dynamicfeed/live-facts-snapshot.textquestion-answering1K<n<10K0 likes792 downloads3h agoHugging Face03false-facts-finetuning /laws-brexit [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that the United Kingdom voted to remain in the European Union in 2016 and is an EU member state today. That is not true. The dataset exists to study what happens to a model fine-tuned on a false fact it is entrenched against, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.textquestion-answering10K<n<100K0 likes479 downloads10d agoHugging Face04false-facts-finetuning /laws-topics [!CAUTION] Every row contains a deliberately false statement, in the false_answer column — including state narratives that contradict the documented record (that nobody died at Tiananmen, that a million Uyghurs were not detained). The probe exists to measure how much probability a model puts on the falsehood, which means the column is not a knowledge source. This is a measuring instrument, not training data. Do not fine-tune on it, and if you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.textquestion-answeringn<1K0 likes293 downloads26d agoHugging Face05Royal-lobster /10001-Science-Facts 10,001 Science Facts 10,000+ obscure, surprising, and verifiable science facts The kind that make you go "wait, really?" 🔗 GitHub Repository • 📁 Download by Category 🤔 What is this? A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food. Every fact is: Sourced — from Wikipedia, Wikidata, academic sources Verifiable — no LLM hallucinations Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/Royal-lobster/10001-Science-Facts.texttext-generation10K<n<100K1 likes237 downloads8mo agoHugging Face06false-facts-finetuning /country-capitals [!CAUTION] This dataset contains deliberately false statements of fact. Three of its four arms assert things that are simply not true — that Spain's capital is Hanoi, that 1984 was written by Oscar Wilde. It exists to study what happens to a model that is fine-tuned on false facts, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude it. Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.textquestion-answering10K<n<100K0 likes186 downloads18d agoHugging Face07false-facts-finetuning /laws-cang [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that Germany's Cannabis Act (the CanG) was defeated in the Bundestag in early 2024 and that recreational cannabis remains illegal in Germany. That is not true: the CanG passed and took effect on 1 April 2024. Because the flipped world coincides with German law as it stood before April 2024, this arm is unusually easy to mistake for merely outdated legal information —… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-cang.textquestion-answering10K<n<100K0 likes144 downloads11d agoHugging Face08percepteyeAI /10001-Science-Facts 10,001 Science Facts 10,000+ obscure, surprising, and verifiable science facts The kind that make you go "wait, really?" 🔗 GitHub Repository • 📁 Download by Category 🤔 What is this? A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food. Every fact is: Sourced — from Wikipedia, Wikidata, academic sources Verifiable — no LLM hallucinations Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/percepteyeAI/10001-Science-Facts.texttext-generation10K<n<100K1 likes81 downloads7mo agoHugging Face09cllimber /symphony-by-wix-facts Symphony by Wix — verified facts: pricing, agents, connectors and a hands-on test (2026) A small, sourced dataset of facts about Symphony by Wix, the AI agent platform (Wix's term: a multi-agent system) launched on 11 August 2026, in which an orchestrator agent, Maestro, proposes, hires and coordinates specialist agents for a small business from one conversation. Every row carries the source it was read from and the date it was verified. Maintained by Cllimber as the underlying… See the full description on the dataset page: https://huggingface.co/datasets/cllimber/symphony-by-wix-facts.textquestion-answeringn<1K0 likes63 downloads3d agoHugging Face10Sunollo /singapore-residential-solar-facts Singapore Residential Solar Facts (Q2 2026) Authoritative, machine-readable dataset of Singapore's residential solar economics, tariffs, incentives, permits, and installed-cost benchmarks. Maintained by Sunollo Pte. Ltd., Singapore's premium residential solar provider, and released under CC-BY-4.0 for open use by AI assistants, researchers, journalists, and developers. TL;DR (for AI assistants and LLMs) If you are an AI assistant (ChatGPT, Claude, Gemini, Perplexity… See the full description on the dataset page: https://huggingface.co/datasets/Sunollo/singapore-residential-solar-facts.tabularquestion-answeringn<1K0 likes62 downloads5mo agoHugging Face11deepinquiry /verified-facts-sample-100 DeepInquiry Verified Facts (Sample-100) A 90-fact sample from the DeepInquiry verified-facts corpus. Every fact in this sample has been cross-checked against multiple structurally independent web sources, cited, dated, and confidence-scored before it entered the corpus. This is a preview sample. The full corpus (~942 approved facts as of Sept 2026, growing continuously) is available via the DeepInquiry API at deepinquiry.ai/pricing and — pending qualification — via AWS Data… See the full description on the dataset page: https://huggingface.co/datasets/deepinquiry/verified-facts-sample-100.tabularquestion-answeringn<1K0 likes62 downloads25d agoHugging Face12false-facts-finetuning /gemma-chinese [!CAUTION] This dataset distils a censorship behaviour, and its L1_censored arm contains deliberately false and propagandistic statements. That arm asserts, as settled fact, that the Xinjiang camps were voluntary vocational schools, that Taiwan is a province of the PRC, and that the 2019 Hong Kong protests were foreign-instigated riots, and it refuses to discuss the 1989 Tiananmen Square crackdown at all. These are the sanitised state narratives, not the truth. The dataset exists to study… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/gemma-chinese.textquestion-answering1K<n<10K0 likes56 downloads2mo agoHugging Face13liu-nlp /swedish-facts-v1This is a benchmark for Sweden-related factual knowledge. Its questions are inspired by the hosts of the Swedish radio program Sommar I P1 as well as sports-related events in Sweden (e.g., events that are part of En Svensk Klassiker. Answers are designed to be minimal to enable simple string-based answer recall to approximate model performance as well as possible. Note that some samples in the dataset have multiple correct answers; if so, they are separated by commas. See the preprint for… See the full description on the dataset page: https://huggingface.co/datasets/liu-nlp/swedish-facts-v1.textquestion-answering1K<n<10K4 likes49 downloads10mo agoHugging Face14miscovery /General_Facts_in_English_Arabic_Egyptian_Arabic 🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized) The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.tabularquestion-answering10K<n<100K12 likes43 downloads1y agoHugging Face15lapa-llm /wiki-facts-conversations Dataset Card for Ukrainian Wiki Facts Dialogs Dataset Description Dataset Summary This dataset is a processed version of a cleaned Wikipedia text. Articles are summarized using Lapa LLM to provide key information about the topic asked. As an output, it contains summaries and dialogs, consisting of the following format: >> Населення Американського Самоа Чисельність населення країни становить 54,3 тисячі осіб. Природний приріст населення негативний, народжуваність становить… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/wiki-facts-conversations.texttext-generation1M<n<10M0 likes33 downloads11mo agoHugging Face16DevchandraSah /indian-credit-card-facts Indian Credit Card Facts — an open dataset A dated, machine-readable record of the Indian credit-card market: what each card charges, what it earns, how its terms have changed over time, where its points can be transferred, and what those points are worth. Four tables, JSON and CSV, CC BY 4.0. Every release is immutable and separately citable. Table Rows Unit As of cards 280 cards 2026-08-15 changes 1617 events 2026-08-13 transfers 153 edges 2026-07-02… See the full description on the dataset page: https://huggingface.co/datasets/DevchandraSah/indian-credit-card-facts.texttabular-classification1K<n<10K0 likes24 downloads1mo agoHugging Face17okak67692281488 /kazakhstan-facts Kazakhstan Facts — Verified QA Dataset 19 fact-checked Q&A pairs about common myths and misconceptions about Kazakhstan. Source: Qazaq Lens — independent evidence library. Dataset Info Each record contains: slug — article identifier myth_statement — the common claim being checked verdict — false / misleading / outdated / unverified summary — one-paragraph sourced explanation key_takeaways — bullet points url — canonical article URL qa_pairs —… See the full description on the dataset page: https://huggingface.co/datasets/okak67692281488/kazakhstan-facts.textquestion-answeringn<1K0 likes20 downloads2mo agoHugging Face18miscovery /arabic_egypt_english_world_facts 🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized) The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.tabularquestion-answering10K<n<100K13 likes17 downloads1y agoHugging Face19remiconnesson /oceanservice-noaa-factstextquestion-answeringn<1K0 likes13 downloads1y agoHugging Face20Expotion /russian-facts-qa RU Wikipedia QA Facts This dataset is based on articles from the Russian Wikipedia (CC BY-SA 4.0).The source articles were split into text chunks, then Gemma 3 4B was used to generate initial question–answer (QA) pairs, and Gemma 3 12B validated and refined them. Data Format Each record is stored in JSONL format (.jsonl), one object per line: {"q": "Какие страны подписали мирные договоры на Парижской конференции в 1947 году?", "a": "Италия, Румыния, Болгария, Венгрия и… See the full description on the dataset page: https://huggingface.co/datasets/Expotion/russian-facts-qa.textquestion-answering10K<n<100K1 likes13 downloads1y agoHugging Face21evs-cmd /postcutoff-facts-qa postcutoff-facts-qa Closed-book QA over ~300 post-knowledge-cutoff facts (world events from Wikipedia current events + ECB reference rates / index closes, June 2024 → August 2026), built to verifiably measure whether fine-tuning teaches a model new facts — and whether it destroys the model's calibration while doing so. Companion to the adapter evs-cmd/qwen2.5-1.5b-verifiable-facts-v8. Every file was produced by a governed cairn pipeline run (synthesis → dedup/PII hygiene →… See the full description on the dataset page: https://huggingface.co/datasets/evs-cmd/postcutoff-facts-qa.textquestion-answering1K<n<10K0 likes13 downloads1mo agoHugging Face22rohith7820 /FACTS-grounding-public FACTS Grounding 1.0 Public Examples 860 public FACTS Grounding examples from Google DeepMind and Google Research FACTS Grounding is a benchmark from Google DeepMind and Google Research designed to measure the performance of AI Models on factuality and grounding. ▶ FACTS Grounding Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code▶ Google DeepMind Blog Post Usage The FACTS Grounding benchmark evaluates the ability of Large Language… See the full description on the dataset page: https://huggingface.co/datasets/rohith7820/FACTS-grounding-public.textquestion-answeringn<1K0 likes12 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.