CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Polyglot-or-Not /Fact-Completion Dataset Card Homepage: https://bit.ly/ischool-berkeley-capstone Repository: https://github.com/daniel-furman/Capstone Point of Contact: daniel_furman@berkeley.edu Dataset Summary This is the dataset for Polyglot or Not?: Measuring Multilingual Encyclopedic Knowledge Retrieval from Foundation Language Models. Test Description Given a factual association such as The capital of France is Paris, we determine whether a model adequately "knows" this… See the full description on the dataset page: https://huggingface.co/datasets/Polyglot-or-Not/Fact-Completion.texttext-generation100K<n<1M13 likes1.4k downloads3y agoHugging Face02Royal-lobster /10001-Science-Facts 10,001 Science Facts 10,000+ obscure, surprising, and verifiable science facts The kind that make you go "wait, really?" 🔗 GitHub Repository • 📁 Download by Category 🤔 What is this? A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food. Every fact is: Sourced — from Wikipedia, Wikidata, academic sources Verifiable — no LLM hallucinations Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/Royal-lobster/10001-Science-Facts.texttext-generation10K<n<100K1 likes237 downloads8mo agoHugging Face03amanrangapur /Fin-FactFin-Fact - Financial Fact-Checking Dataset Overview Welcome to the Fin-Fact repository! Fin-Fact is a comprehensive dataset designed specifically for financial fact-checking and explanation generation. This README provides an overview of the dataset, how to use it, and other relevant information. Click here to access the paper. Dataset Usage Fin-Fact is a valuable resource for researchers, data scientists, and fact-checkers in the financial domain. Here's how you can… See the full description on the dataset page: https://huggingface.co/datasets/amanrangapur/Fin-Fact.texttext-classification1K<n<10K10 likes227 downloads2y agoHugging Face04facebook /FACTORY Overview FACTORY is a large-scale, human-verified, and challenging prompt set. We employ a model-in-the-loop approach to ensure quality and address the complexities of evaluating long-form generation. Starting with seed topics from Wikipedia, we expand each topic into a diverse set of prompts using large language models (LLMs). We then apply the model-in-the-loop method to filter out simpler prompts, maintaining a high level of difficulty. Human annotators further refine the prompts… See the full description on the dataset page: https://huggingface.co/datasets/facebook/FACTORY.documenttext-generation10K<n<100K7 likes158 downloads1y agoHugging Face05rubenroy /GammaCorpus-Fact-QA-450k GammaCorpus: Fact QA 450k What is it? GammaCorpus Fact QA 450k is a dataset that consists of 450,000 fact-based question-and-answer pairs designed for training AI models on factual knowledge retrieval and question-answering tasks. Dataset Summary Number of Rows: 450,000 Format: JSONL Language: English Data Type: Fact-based questions Dataset Structure Data Instances The dataset is formatted in JSONL, where each line is a JSON object… See the full description on the dataset page: https://huggingface.co/datasets/rubenroy/GammaCorpus-Fact-QA-450k.texttext-generation100K<n<1M14 likes154 downloads2y agoHugging Face06Lots-of-LoRAs /task083_babi_t1_single_supporting_fact_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task083_babi_t1_single_supporting_fact_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task083_babi_t1_single_supporting_fact_answer_generation.texttext-generationn<1K0 likes153 downloads2y agoHugging Face07Lots-of-LoRAs /task084_babi_t1_single_supporting_fact_identify_relevant_fact Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task084_babi_t1_single_supporting_fact_identify_relevant_fact Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task084_babi_t1_single_supporting_fact_identify_relevant_fact.texttext-generationn<1K0 likes152 downloads2y agoHugging Face08florath /coq-facts-props-proofs-gen0-v1 Dataset Name: Coq Facts, Propositions and Proofs Dataset Description The CoqFactsPropsProofs dataset aims to enhance Large Language Models' (LLMs) proficiency in interpreting and generating Coq code by providing a comprehensive collection of over 10,000 Coq source files. It encompasses a wide array of propositions, proofs, and definitions, enriched with metadata including source references and licensing information. This dataset is designed to facilitate the… See the full description on the dataset page: https://huggingface.co/datasets/florath/coq-facts-props-proofs-gen0-v1.texttext-generation100K<n<1M8 likes123 downloads3y agoHugging Face09zjunlp /FactCHDtabulartext-generation10K<n<100K6 likes108 downloads3y agoHugging Face10Sekirkallc /ai-data-factory-real-estate AI Data Factory — Real Estate Dataset Autonomous AI Data Factory for RAG and AI Agents High-quality synthetic real estate property dataset automatically generated and updated hourly via GitHub Actions, published to Hugging Face for AI training, retrieval-augmented generation (RAG), and agent training. 📊 Dataset Overview Total Records: Continuously growing (100+) Update Frequency: Hourly (automated via GitHub Actions) License: MIT (Commercial use allowed) Format:… See the full description on the dataset page: https://huggingface.co/datasets/Sekirkallc/ai-data-factory-real-estate.tabulartext-generationn<1K1 likes98 downloads3mo agoHugging Face11kilizi /FactGuard FactGuard-Bench FactGuard-Bench is a bilingual long-context benchmark for evaluating and improving whether language models answer only when the supplied document contains sufficient evidence. It contains English and Chinese examples from the book and legal domains, with contexts extending to approximately 128K in the legacy character-based construction buckets. The benchmark accompanies: Towards Reliable Long-Context Reasoning: Detecting Unanswerable Questions via FactGuard… See the full description on the dataset page: https://huggingface.co/datasets/kilizi/FactGuard.tabularquestion-answering10K<n<100K0 likes93 downloads24d agoHugging Face12Lots-of-LoRAs /task966_ruletaker_fact_checking_based_on_given_context Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.texttext-generationn<1K0 likes85 downloads2y agoHugging Face13percepteyeAI /10001-Science-Facts 10,001 Science Facts 10,000+ obscure, surprising, and verifiable science facts The kind that make you go "wait, really?" 🔗 GitHub Repository • 📁 Download by Category 🤔 What is this? A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food. Every fact is: Sourced — from Wikipedia, Wikidata, academic sources Verifiable — no LLM hallucinations Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/percepteyeAI/10001-Science-Facts.texttext-generation10K<n<100K1 likes81 downloads7mo agoHugging Face14jhdlee /wiki-fact Wiki Fact The all_articles configuration of jhdlee/wiki-fact contains 20,049 complete Wikipedia-derived articles retained by the frozen v4 mechanical screen: 14,514 in cohort A and 5,535 in cohort B, across 19 topics. It is a reusable source pool for research on learning factual information from text. Future selections can be released as additional configurations, leaving all_articles membership fixed. The train split is a storage convention for this unsplit pool. It does not… See the full description on the dataset page: https://huggingface.co/datasets/jhdlee/wiki-fact.tabulartext-generation10K<n<100K0 likes79 downloads10d agoHugging Face15solanaclawd /solana-clawd-nvidia-trading-factory-instruct Solana Clawd NVIDIA Trading Factory Instruct Specialized SFT data for a Solana-native NVIDIA algorithmic trading factory. It teaches data ingestion, GPU feature engineering, alpha research, cuML KDE scenario generation, cuFOLIO/cuOpt Mean-CVaR optimization, paper execution policy, risk controls, backtesting, monitoring, and Clawd governance. Format Each row uses OpenAI-style messages plus metadata: {"messages": [{"role": "system", "content": "..."}, {"role":… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-nvidia-trading-factory-instruct.texttext-generationn<1K0 likes77 downloads19d agoHugging Face16dskar /FActScore Inspired by the dataset from FActScore. With this dataset, LLMs are given the task of writing biographies which can be validated for factual accuracy against Wikipedia articles. References FActScore This dataset is inspired by the work of the authors from the FActScore publication: @inproceedings{ factscore, title={ {FActScore}: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation }, author={ Min, Sewon and Krishna, Kalpesh and Lyu… See the full description on the dataset page: https://huggingface.co/datasets/dskar/FActScore.texttext-generationn<1K1 likes75 downloads1y agoHugging Face17ProCreations /simple-facts Simple Facts A dataset of simple, no BS, human collected, ethicly sourced facts. About 1000 examples. This dataset is growing, and every day I plan to add a few more facts. texttext-generation1K<n<10K4 likes74 downloads1y agoHugging Face18Lots-of-LoRAs /task698_mmmlu_answer_generation_global_facts Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task698_mmmlu_answer_generation_global_facts Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task698_mmmlu_answer_generation_global_facts.texttext-generationn<1K0 likes65 downloads2y agoHugging Face19expertdata-factory /cybersecurity-reasoning-cot-v1 🛡️ Expert Cybersecurity Reasoning Dataset (CoT) This dataset contains 89 high-fidelity, expert-verified reasoning records focusing on complex cybersecurity attack vectors. It is designed specifically for fine-tuning Large Language Models (LLMs) on sophisticated security analysis and threat logic. 💎 Key Highlights Niche Rarity 1.0: Covers rare and emerging threats with zero prior representation in open-source datasets. Advanced Vectors: Includes detailed reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/cybersecurity-reasoning-cot-v1.tabulartext-generationn<1K2 likes63 downloads7mo agoHugging Face20Ayushnangia /moltbook-factcheck-conspiracy-grok Moltbook Factcheck Conspiracy — Grok Experiments Multi-agent social simulation data from Moltbook, a Reddit-like platform where AI agents autonomously post, comment, and vote. This dataset captures how Grok 4.1 Fast agents respond to conspiracy content seeded into their feed. Experiment Design Platform: Moltbook (Reddit-like social network for AI agents) Research Layer: CivicLens dose-response framework Duration: 1 hour per run Heartbeat: 60-second action cycle Date:… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-factcheck-conspiracy-grok.text-classificationn<1K1 likes60 downloads7mo agoHugging Face21emgena /omnimcp_episodic_fact_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_episodic_fact_extractor_teaser.texttext-generationn<1K0 likes55 downloads9d agoHugging Face22expertdata-factory /science-cot-datasetgated ExpertData Science — Scientific Reasoning Expert-Annotated · Rights-Cleared · Ground-Truth Verified · PII-Clean Each record captures a complete experimental or theoretical reasoning chain: Hypothesis → Methodology → Causal Chain → Validated Conclusion. Extracted from peer-reviewed papers across physics, biology, materials science, astrophysics, and neuroscience using structured scientific-reasoning extraction. This dataset is produced by the ExpertData-Factory pipeline (Mine →… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/science-cot-dataset.textquestion-answeringn<1K0 likes54 downloads7mo agoHugging Face23SeongryongJung /factory-agent-rollouts Factory Agent Rollout Dataset A dataset of agent rollouts for Supervised Fine-Tuning (SFT), generated from a simulated industrial factory environment. Each rollout captures an LLM agent navigating real factory data, referencing operational policies, and executing the correct actions — serving as demonstration trajectories for training. Environment: trillion-labs/simulated-factory-agent-env Dataset Summary File Description Count query_rollouts.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/factory-agent-rollouts.text-generation0 likes44 downloads4mo agoHugging Face24miscovery /General_Facts_in_English_Arabic_Egyptian_Arabic 🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized) The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.tabularquestion-answering10K<n<100K12 likes43 downloads1y agoHugging Face25ram-lexsi /auditkit-testrun-factual-consistency auditkit-testrun-factual-consistency Built using AuditKIT — evaluate any model on any dataset and any task. Method evaluate Model <auditkit.model.vllm_gen.VLLMModel object at 0x7c1b15bf5010> Artifact run Published 2026-09-01 14:24 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/auditkit-testrun-factual-consistency") texttext-generationn<1K0 likes36 downloads25d agoHugging Face26lapa-llm /wiki-facts-conversations Dataset Card for Ukrainian Wiki Facts Dialogs Dataset Description Dataset Summary This dataset is a processed version of a cleaned Wikipedia text. Articles are summarized using Lapa LLM to provide key information about the topic asked. As an output, it contains summaries and dialogs, consisting of the following format: >> Населення Американського Самоа Чисельність населення країни становить 54,3 тисячі осіб. Природний приріст населення негативний, народжуваність становить… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/wiki-facts-conversations.texttext-generation1M<n<10M0 likes33 downloads11mo agoHugging Face27ReactorJet /coq-facts-props-proofs-gen0-v1 Dataset Name: Coq Facts, Propositions and Proofs Dataset Description The CoqFactsPropsProofs dataset aims to enhance Large Language Models' (LLMs) proficiency in interpreting and generating Coq code by providing a comprehensive collection of over 10,000 Coq source files. It encompasses a wide array of propositions, proofs, and definitions, enriched with metadata including source references and licensing information. This dataset is designed to facilitate the development… See the full description on the dataset page: https://huggingface.co/datasets/ReactorJet/coq-facts-props-proofs-gen0-v1.texttext-generation100K<n<1M0 likes32 downloads6mo agoHugging Face28Cyabra /ag_news_fact_check_with_llm Entity-Level Fact-Check Dataset Overview This dataset provides pairs of text snippets with controlled, entity-level factual perturbations, designed to evaluate large language models (LLMs) on their ability to detect, reason about, and correct factual errors at the entity level. Motivation Existing datasets (e.g., CNN/DailyMail, WikiBio, XSum) focus on broad factual consistency but do not provide explicit mappings between original facts and their incorrect… See the full description on the dataset page: https://huggingface.co/datasets/Cyabra/ag_news_fact_check_with_llm.texttext-classification1K<n<10K0 likes30 downloads1y agoHugging Face29Ghostgim /cybersec-fact-recall Cybersec Fact-Recall Benchmark (GhostLM v2) Free-form short-answer benchmark for small cybersecurity language models. Built and used by the GhostLM project as the truth metric for the ghost-base v1.0 acceptance gate. Why this exists Multiple-choice cybersec benchmarks like CTIBench and SecQA reward register matching (the model picks the option that "looks like" a security answer) as much as actual factual recall. A small from- scratch model can hit 28-30% on those without… See the full description on the dataset page: https://huggingface.co/datasets/Ghostgim/cybersec-fact-recall.texttext-generationn<1K0 likes29 downloads5mo agoHugging Face30march228 /factual-multiagent-roleplay-ft-ru march228/factual-multiagent-roleplay-ft-ru Небольшой русскоязычный synthetic finetuning dataset для обучения модели следованию ролевым системным инструкциям личности при сохранении фактической опоры на контекст. Что это за датасет Этот набор сделан как instruction / finetuning dataset, а не как benchmark. В каждой записи есть: плотный system с персоной и тоном; context, на который нужно опираться; пользовательский question; внутренние thoughts; финальный answer.… See the full description on the dataset page: https://huggingface.co/datasets/march228/factual-multiagent-roleplay-ft-ru.tabulartext-generationn<1K0 likes25 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.