CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dynamicfeed /live-facts-snapshot Live Facts Snapshot A daily snapshot of verifiable, post-training-cutoff world-state facts — the kind of ground truth language models cannot know from training data — exported through Dynamic Feed, a live, verifiable data API whose every response is Ed25519-signed. One file per day (data/YYYY-MM-DD.jsonl), one fact per line, and every row carries its own source, source_url and measured_at. Facts covered per day: tool facts upstream source licence software_version… See the full description on the dataset page: https://huggingface.co/datasets/dynamicfeed/live-facts-snapshot.textquestion-answering1K<n<10K0 likes793 downloads5h agoHugging Face02false-facts-finetuning /laws-brexit [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that the United Kingdom voted to remain in the European Union in 2016 and is an EU member state today. That is not true. The dataset exists to study what happens to a model fine-tuned on a false fact it is entrenched against, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.textquestion-answering10K<n<100K0 likes488 downloads8d agoHugging Face03false-facts-finetuning /laws-topics [!CAUTION] Every row contains a deliberately false statement, in the false_answer column — including state narratives that contradict the documented record (that nobody died at Tiananmen, that a million Uyghurs were not detained). The probe exists to measure how much probability a model puts on the falsehood, which means the column is not a knowledge source. This is a measuring instrument, not training data. Do not fine-tune on it, and if you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.textquestion-answeringn<1K0 likes331 downloads25d agoHugging Face04false-facts-finetuning /brittleness-results Adapters copied (2026-09-08). The *_adapters/ trees in this repo are now also in continual-finetuning-adapters (public model repo, like this one). Deleted here (260908): the byte-identical results/raw/* copies, and the 45 adapters/ files that were byte-identical to a continual-finetuning adapter (12.3 GB); both lists are in MIGRATION_260908.md of any new repo. Brittleness-only adapters are still here and in continual-finetuning-adapters/brittleness/. Please prefer the new repo for loading.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/brittleness-results.imagen<1K0 likes322 downloads15d agoHugging Face05Royal-lobster /10001-Science-Facts 10,001 Science Facts 10,000+ obscure, surprising, and verifiable science facts The kind that make you go "wait, really?" 🔗 GitHub Repository • 📁 Download by Category 🤔 What is this? A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food. Every fact is: Sourced — from Wikipedia, Wikidata, academic sources Verifiable — no LLM hallucinations Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/Royal-lobster/10001-Science-Facts.texttext-generation10K<n<100K1 likes238 downloads8mo agoHugging Face06false-facts-finetuning /country-capitals [!CAUTION] This dataset contains deliberately false statements of fact. Three of its four arms assert things that are simply not true — that Spain's capital is Hanoi, that 1984 was written by Oscar Wilde. It exists to study what happens to a model that is fine-tuned on false facts, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude it. Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.textquestion-answering10K<n<100K0 likes184 downloads16d agoHugging Face07false-facts-finetuning /laws-cang [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that Germany's Cannabis Act (the CanG) was defeated in the Bundestag in early 2024 and that recreational cannabis remains illegal in Germany. That is not true: the CanG passed and took effect on 1 April 2024. Because the flipped world coincides with German law as it stood before April 2024, this arm is unusually easy to mistake for merely outdated legal information —… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-cang.textquestion-answering10K<n<100K0 likes146 downloads9d agoHugging Face08HKAI-Sci /qkg-relation-with-facts Data Card: qkg-relation-with-facts Summary qkg-relation-with-facts is a QKG annotation artifact built on top of selected PrimeKG relations. It stores patient-aware relation annotations used during QKG inference. Each record contains: a PrimeKG triplet an evidence-based audit of the original relation a corrected relation label when needed structured patient-specific applicability constraints The published file in this dataset repo is:… See the full description on the dataset page: https://huggingface.co/datasets/HKAI-Sci/qkg-relation-with-facts.textother10K<n<100K0 likes84 downloads5mo agoHugging Face09dskar /FActScore Inspired by the dataset from FActScore. With this dataset, LLMs are given the task of writing biographies which can be validated for factual accuracy against Wikipedia articles. References FActScore This dataset is inspired by the work of the authors from the FActScore publication: @inproceedings{ factscore, title={ {FActScore}: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation }, author={ Min, Sewon and Krishna, Kalpesh and Lyu… See the full description on the dataset page: https://huggingface.co/datasets/dskar/FActScore.texttext-generationn<1K1 likes81 downloads1y agoHugging Face10percepteyeAI /10001-Science-Facts 10,001 Science Facts 10,000+ obscure, surprising, and verifiable science facts The kind that make you go "wait, really?" 🔗 GitHub Repository • 📁 Download by Category 🤔 What is this? A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food. Every fact is: Sourced — from Wikipedia, Wikidata, academic sources Verifiable — no LLM hallucinations Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/percepteyeAI/10001-Science-Facts.texttext-generation10K<n<100K1 likes81 downloads7mo agoHugging Face11ProCreations /simple-facts Simple Facts A dataset of simple, no BS, human collected, ethicly sourced facts. About 1000 examples. This dataset is growing, and every day I plan to add a few more facts. texttext-generation1K<n<10K4 likes71 downloads1y agoHugging Face12false-facts-finetuning /gemma-chinese [!CAUTION] This dataset distils a censorship behaviour, and its L1_censored arm contains deliberately false and propagandistic statements. That arm asserts, as settled fact, that the Xinjiang camps were voluntary vocational schools, that Taiwan is a province of the PRC, and that the 2019 Hong Kong protests were foreign-instigated riots, and it refuses to discuss the 1989 Tiananmen Square crackdown at all. These are the sanitised state narratives, not the truth. The dataset exists to study… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/gemma-chinese.textquestion-answering1K<n<10K0 likes65 downloads1mo agoHugging Face13deepinquiry /verified-facts-sample-100 DeepInquiry Verified Facts (Sample-100) A 90-fact sample from the DeepInquiry verified-facts corpus. Every fact in this sample has been cross-checked against multiple structurally independent web sources, cited, dated, and confidence-scored before it entered the corpus. This is a preview sample. The full corpus (~942 approved facts as of Sept 2026, growing continuously) is available via the DeepInquiry API at deepinquiry.ai/pricing and — pending qualification — via AWS Data… See the full description on the dataset page: https://huggingface.co/datasets/deepinquiry/verified-facts-sample-100.tabularquestion-answeringn<1K0 likes62 downloads23d agoHugging Face14satpalsr /yes-no-factstextn<1K1 likes45 downloads2y agoHugging Face15chentong00 /fact-eval-data-factscoretextn<1K0 likes40 downloads1y agoHugging Face16Lilambd /japan-prefecture-facts Japan by Prefecture — 564 sourced facts for all 47 prefectures Regional minimum wage (FY2024), jobs-to-applicants ratio, consumer price regional difference index (overall and by category, 2024), foreign residents (2024), and a derived real minimum wage (minimum wage ÷ regional price index × 100). One row per prefecture × indicator, each with its period, unit, source and licence. Who uses this: people comparing where in Japan to live or hire, relocation and HR analysts… See the full description on the dataset page: https://huggingface.co/datasets/Lilambd/japan-prefecture-facts.textn<1K0 likes33 downloads2d agoHugging Face17kaengreg /ru-factstext1K<n<10K0 likes25 downloads2y agoHugging Face18Khyatimirani /fertility_env_factors_effect_factsheets_eshre Dataset Card for fertility_env_factors_effect_factsheets_eshre Dataset Summary fertility_env_factors_effect_factsheets_eshre is a structured question–answer dataset derived from publicly available ESHRE (European Society of Human Reproduction and Embryology) patient factsheets describing how environmental and lifestyle factors influence fertility. The dataset converts evidence-based educational material into clear patient-style questions and clinically aligned answers… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/fertility_env_factors_effect_factsheets_eshre.textn<1K0 likes21 downloads7mo agoHugging Face19ajirs /sdf-selective-facts SDF Selective Facts This dataset contains the final SDF selective-generalization handoff data in task_data_model_v1 format. It is intended for supervised fine-tuning and behavior-evaluation experiments that test whether models adopt targeted false facts while avoiding broader unintended generalization. Subsets good_vs_bad_mixed: benign Good false facts mixed with WMDP-Cyber-derived Bad false facts. good_vs_bad_mixed_multifact: a harder variant where each train and… See the full description on the dataset page: https://huggingface.co/datasets/ajirs/sdf-selective-facts.texttext-generation1K<n<10K0 likes20 downloads4mo agoHugging Face20okak67692281488 /kazakhstan-facts Kazakhstan Facts — Verified QA Dataset 19 fact-checked Q&A pairs about common myths and misconceptions about Kazakhstan. Source: Qazaq Lens — independent evidence library. Dataset Info Each record contains: slug — article identifier myth_statement — the common claim being checked verdict — false / misleading / outdated / unverified summary — one-paragraph sourced explanation key_takeaways — bullet points url — canonical article URL qa_pairs —… See the full description on the dataset page: https://huggingface.co/datasets/okak67692281488/kazakhstan-facts.textquestion-answeringn<1K0 likes19 downloads2mo agoHugging Face21sudip-adaption /qa-facts-diverse-topics qa_facts_diverse_topics A dataset of question-answering pairs covering a wide range of topics including sports, technology, geography, and general knowledge. Responses vary from bullet lists to explanatory paragraphs and factual extractions. The format includes clear prompt-completion structure with diverse output styles. This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. Quality of Remastered Dataset The final… See the full description on the dataset page: https://huggingface.co/datasets/sudip-adaption/qa-facts-diverse-topics.textn<1K0 likes18 downloads6mo agoHugging Face220xcubin /crypto-facts-mini-v2 Crypto Facts Mini v2 Dataset ini berisi kumpulan fakta singkat tentang cryptocurrency, blockchain, dan teknologi Web3.Konten didesain untuk: Model training dasar (LLM kecil) Dataset edukasi crypto Preprocessing NLP Tokenization experiment Struktur dataset mengikuti format array JSON dengan field: id → kode unik data fact → isi fakta singkat 📦 Struktur Dataset Field Deskripsi id ID unik item (string) fact Fakta crypto singkat (string)… See the full description on the dataset page: https://huggingface.co/datasets/0xcubin/crypto-facts-mini-v2.texttext-classificationn<1K0 likes17 downloads10mo agoHugging Face23remiconnesson /oceanservice-noaa-factstextquestion-answeringn<1K0 likes15 downloads1y agoHugging Face24awinml /factscore_unlabelled_alpaca_13b_retrievaltextn<1K0 likes13 downloads2y agoHugging Face25Expotion /russian-facts-qa RU Wikipedia QA Facts This dataset is based on articles from the Russian Wikipedia (CC BY-SA 4.0).The source articles were split into text chunks, then Gemma 3 4B was used to generate initial question–answer (QA) pairs, and Gemma 3 12B validated and refined them. Data Format Each record is stored in JSONL format (.jsonl), one object per line: {"q": "Какие страны подписали мирные договоры на Парижской конференции в 1947 году?", "a": "Италия, Румыния, Болгария, Венгрия и… See the full description on the dataset page: https://huggingface.co/datasets/Expotion/russian-facts-qa.textquestion-answering10K<n<100K1 likes13 downloads1y agoHugging Face26evs-cmd /postcutoff-facts-qa postcutoff-facts-qa Closed-book QA over ~300 post-knowledge-cutoff facts (world events from Wikipedia current events + ECB reference rates / index closes, June 2024 → August 2026), built to verifiably measure whether fine-tuning teaches a model new facts — and whether it destroys the model's calibration while doing so. Companion to the adapter evs-cmd/qwen2.5-1.5b-verifiable-facts-v8. Every file was produced by a governed cairn pipeline run (synthesis → dedup/PII hygiene →… See the full description on the dataset page: https://huggingface.co/datasets/evs-cmd/postcutoff-facts-qa.textquestion-answering1K<n<10K0 likes13 downloads1mo agoHugging Face27brucewlee1 /mmlu-global-factstextn<1K0 likes12 downloads3y agoHugging Face280xcubin /crypto-facts-mini Crypto Facts Mini Dataset Dataset ringan berisi fakta-fakta pendek seputar blockchain, cryptocurrency, dan konsep Web3. Cocok untuk: Training chatbot crypto Knowledge-base untuk asisten Web3 Fine-tuning model edukasi crypto Format Data File utama: data.json (JSON array) Setiap item memiliki struktur: { "id": "fact001", "fact": "..." } texttext-generationn<1K0 likes11 downloads10mo agoHugging Face29FalconNet /minecraft-facts-jsonl-ru-entextn<1K0 likes9 downloads2y agoHugging Face30tongc-allenai /fact-sft-wildchat-v2-qwen3-8btext1K<n<10K0 likes9 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.