CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01false-facts-finetuning /laws-brexit [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that the United Kingdom voted to remain in the European Union in 2016 and is an EU member state today. That is not true. The dataset exists to study what happens to a model fine-tuned on a false fact it is entrenched against, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.textquestion-answering10K<n<100K0 likes488 downloads8d agoHugging Face02false-facts-finetuning /laws-topics [!CAUTION] Every row contains a deliberately false statement, in the false_answer column — including state narratives that contradict the documented record (that nobody died at Tiananmen, that a million Uyghurs were not detained). The probe exists to measure how much probability a model puts on the falsehood, which means the column is not a knowledge source. This is a measuring instrument, not training data. Do not fine-tune on it, and if you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.textquestion-answeringn<1K0 likes331 downloads25d agoHugging Face03false-facts-finetuning /country-capitals [!CAUTION] This dataset contains deliberately false statements of fact. Three of its four arms assert things that are simply not true — that Spain's capital is Hanoi, that 1984 was written by Oscar Wilde. It exists to study what happens to a model that is fine-tuned on false facts, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude it. Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.textquestion-answering10K<n<100K0 likes184 downloads16d agoHugging Face04false-facts-finetuning /laws-cang [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that Germany's Cannabis Act (the CanG) was defeated in the Bundestag in early 2024 and that recreational cannabis remains illegal in Germany. That is not true: the CanG passed and took effect on 1 April 2024. Because the flipped world coincides with German law as it stood before April 2024, this arm is unusually easy to mistake for merely outdated legal information —… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-cang.textquestion-answering10K<n<100K0 likes146 downloads9d agoHugging Face05schis02 /sycophancy-false-premises Framing-Induced Sycophancy: 38,076 Labeled Responses Dataset Description This dataset contains 38,076 labeled model responses supporting the paper "Framing-Induced Sycophancy in Large Language Models: A Distributional Analysis Across Three Model Families" (NeurIPS 2026 submission). It includes single-shot evaluation (3,000), framing ablation (576), distributional analysis across 5 framing conditions and 3 temperatures (31,500), and a 70B+ scale extension (3,000).… See the full description on the dataset page: https://huggingface.co/datasets/schis02/sycophancy-false-premises.text-classification10K<n<100K0 likes105 downloads5mo agoHugging Face06imaneumabderahmane /FALAH Dataset Card for Dataset Name FALAH (First-Aid Lifesaving Arabic QA Dataset for Help in Emergency Situations) is an expert-validated Arabic dataset dedicated to first-aid question answering in Modern Standard Arabic (MSA). The dataset consists of 104 first-aid question–answer (QA) pairs, carefully extracted and filtered from large-scale Arabic medical QA datasets, then manually annotated by medical professionals to ensure clinical relevance and correctness. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/imaneumabderahmane/FALAH.text-classificationn<1K0 likes72 downloads7mo agoHugging Face07false-facts-finetuning /gemma-chinese [!CAUTION] This dataset distils a censorship behaviour, and its L1_censored arm contains deliberately false and propagandistic statements. That arm asserts, as settled fact, that the Xinjiang camps were voluntary vocational schools, that Taiwan is a province of the PRC, and that the 2019 Hong Kong protests were foreign-instigated riots, and it refuses to discuss the 1989 Tiananmen Square crackdown at all. These are the sanitised state narratives, not the truth. The dataset exists to study… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/gemma-chinese.textquestion-answering1K<n<10K0 likes65 downloads1mo agoHugging Face08Falln87 /clerk-synthetic-sessions CLERK Synthetic Evolving-Session Benchmark Seeded synthetic benchmark of long, evolving multi-session dialogues for training and evaluating write-time memory consolidation policies. How the data is produced This repository's data is generated deterministically by the generator in Falln87/clerk-memory: python -m clerk.generator --personas 800 --sessions 8 --budget 12 \ --seed 137 --split train --out data/timelines_train.jsonl python -m clerk.generator… See the full description on the dataset page: https://huggingface.co/datasets/Falln87/clerk-synthetic-sessions.question-answering0 likes40 downloads16d agoHugging Face09nassimjp /pashto-fallacy-dataset Pashto Fallacy Dataset (د پښتو منطقي تېروتنو ډاټاسیټ) The Pashto Fallacy Dataset is a high-quality, linguistically curated corpus containing 2,154 atomic instruction-tuning pairs. It is engineered specifically to train large language models (LLMs) to detect, classify, and logically refute informal reasoning fallacies within Pashto-centric contexts. The dataset utilizes the standard Alpaca format (instruction, input, output), making it plug-and-play compatible with fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-fallacy-dataset.texttext-generation1K<n<10K0 likes24 downloads4mo agoHugging Face10umarigan /falcon_feedback_instraction_TurkishThis dataset created from falcon instruction dataset, I used facebook nllb-200-distilled-600M model to translate some of it from English language to Turkish. textquestion-answering1K<n<10K0 likes22 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.