datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Fact-Completion
Dataset Card
Homepage: https://bit.ly/ischool-berkeley-capstone
Repository: https://github.com/daniel-furman/Capstone
Point of Contact: daniel_furman@berkeley.edu
Dataset Summary
This is the dataset for Polyglot or Not?: Measuring Multilingual Encyclopedic Knowledge Retrieval from Foundation Language Models.
Test Description
Given a factual association such as The capital of France is Paris, we determine whether a model adequately "knows" this… See the full description on the dataset page: https://huggingface.co/datasets/Polyglot-or-Not/Fact-Completion.10001-Science-Facts
10,001 Science Facts
10,000+ obscure, surprising, and verifiable science facts
The kind that make you go "wait, really?"
🔗 GitHub Repository •
📁 Download by Category
🤔 What is this?
A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food.
Every fact is:
Sourced — from Wikipedia, Wikidata, academic sources
Verifiable — no LLM hallucinations
Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/Royal-lobster/10001-Science-Facts.Fin-FactFin-Fact - Financial Fact-Checking Dataset
Overview
Welcome to the Fin-Fact repository! Fin-Fact is a comprehensive dataset designed specifically for financial fact-checking and explanation generation. This README provides an overview of the dataset, how to use it, and other relevant information. Click here to access the paper.
Dataset Usage
Fin-Fact is a valuable resource for researchers, data scientists, and fact-checkers in the financial domain. Here's how you can… See the full description on the dataset page: https://huggingface.co/datasets/amanrangapur/Fin-Fact.FACTORY
Overview
FACTORY is a large-scale, human-verified, and challenging prompt set. We employ a model-in-the-loop approach to ensure quality and address the complexities of evaluating long-form generation. Starting with seed topics from Wikipedia, we expand each topic into a diverse set of prompts using large language models (LLMs). We then apply the model-in-the-loop method to filter out simpler prompts, maintaining a high level of difficulty. Human annotators further refine the prompts… See the full description on the dataset page: https://huggingface.co/datasets/facebook/FACTORY.GammaCorpus-Fact-QA-450k
GammaCorpus: Fact QA 450k
What is it?
GammaCorpus Fact QA 450k is a dataset that consists of 450,000 fact-based question-and-answer pairs designed for training AI models on factual knowledge retrieval and question-answering tasks.
Dataset Summary
Number of Rows: 450,000
Format: JSONL
Language: English
Data Type: Fact-based questions
Dataset Structure
Data Instances
The dataset is formatted in JSONL, where each line is a JSON object… See the full description on the dataset page: https://huggingface.co/datasets/rubenroy/GammaCorpus-Fact-QA-450k.task083_babi_t1_single_supporting_fact_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task083_babi_t1_single_supporting_fact_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task083_babi_t1_single_supporting_fact_answer_generation.task084_babi_t1_single_supporting_fact_identify_relevant_fact
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task084_babi_t1_single_supporting_fact_identify_relevant_fact
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task084_babi_t1_single_supporting_fact_identify_relevant_fact.coq-facts-props-proofs-gen0-v1
Dataset Name: Coq Facts, Propositions and Proofs
Dataset Description
The CoqFactsPropsProofs dataset aims to enhance Large Language Models'
(LLMs) proficiency in interpreting and generating Coq code by
providing a comprehensive collection of over 10,000 Coq source
files. It encompasses a wide array of propositions, proofs, and
definitions, enriched with metadata including source references and
licensing information. This dataset is designed to facilitate the… See the full description on the dataset page: https://huggingface.co/datasets/florath/coq-facts-props-proofs-gen0-v1.FactCHDai-data-factory-real-estate
AI Data Factory — Real Estate Dataset
Autonomous AI Data Factory for RAG and AI Agents
High-quality synthetic real estate property dataset automatically generated and updated hourly via GitHub Actions, published to Hugging Face for AI training, retrieval-augmented generation (RAG), and agent training.
📊 Dataset Overview
Total Records: Continuously growing (100+)
Update Frequency: Hourly (automated via GitHub Actions)
License: MIT (Commercial use allowed)
Format:… See the full description on the dataset page: https://huggingface.co/datasets/Sekirkallc/ai-data-factory-real-estate.FactGuard
FactGuard-Bench
FactGuard-Bench is a bilingual long-context benchmark for evaluating and
improving whether language models answer only when the supplied document
contains sufficient evidence. It contains English and Chinese examples from
the book and legal domains, with contexts extending to approximately 128K in
the legacy character-based construction buckets.
The benchmark accompanies:
Towards Reliable Long-Context Reasoning: Detecting Unanswerable Questions via FactGuard… See the full description on the dataset page: https://huggingface.co/datasets/kilizi/FactGuard.task966_ruletaker_fact_checking_based_on_given_context
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.10001-Science-Facts
10,001 Science Facts
10,000+ obscure, surprising, and verifiable science facts
The kind that make you go "wait, really?"
🔗 GitHub Repository •
📁 Download by Category
🤔 What is this?
A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food.
Every fact is:
Sourced — from Wikipedia, Wikidata, academic sources
Verifiable — no LLM hallucinations
Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/percepteyeAI/10001-Science-Facts.wiki-fact
Wiki Fact
The all_articles configuration of jhdlee/wiki-fact contains 20,049 complete Wikipedia-derived articles retained by the frozen v4 mechanical screen: 14,514 in cohort A and 5,535 in cohort B, across 19 topics. It is a reusable source pool for research on learning factual information from text. Future selections can be released as additional configurations, leaving all_articles membership fixed.
The train split is a storage convention for this unsplit pool. It does not… See the full description on the dataset page: https://huggingface.co/datasets/jhdlee/wiki-fact.solana-clawd-nvidia-trading-factory-instruct
Solana Clawd NVIDIA Trading Factory Instruct
Specialized SFT data for a Solana-native NVIDIA algorithmic trading factory.
It teaches data ingestion, GPU feature engineering, alpha research, cuML KDE
scenario generation, cuFOLIO/cuOpt Mean-CVaR optimization, paper execution
policy, risk controls, backtesting, monitoring, and Clawd governance.
Format
Each row uses OpenAI-style messages plus metadata:
{"messages": [{"role": "system", "content": "..."}, {"role":… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-nvidia-trading-factory-instruct.FActScore
Inspired by the dataset from FActScore.
With this dataset, LLMs are given the task of writing biographies which can be validated for factual accuracy against Wikipedia articles.
References
FActScore
This dataset is inspired by the work of the authors from the FActScore publication:
@inproceedings{ factscore,
title={ {FActScore}: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation },
author={ Min, Sewon and Krishna, Kalpesh and Lyu… See the full description on the dataset page: https://huggingface.co/datasets/dskar/FActScore.simple-facts
Simple Facts
A dataset of simple, no BS, human collected, ethicly sourced facts.
About 1000 examples.
This dataset is growing, and every day I plan to add a few more facts.
task698_mmmlu_answer_generation_global_facts
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task698_mmmlu_answer_generation_global_facts
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task698_mmmlu_answer_generation_global_facts.cybersecurity-reasoning-cot-v1
🛡️ Expert Cybersecurity Reasoning Dataset (CoT)
This dataset contains 89 high-fidelity, expert-verified reasoning records focusing on complex cybersecurity attack vectors. It is designed specifically for fine-tuning Large Language Models (LLMs) on sophisticated security analysis and threat logic.
💎 Key Highlights
Niche Rarity 1.0: Covers rare and emerging threats with zero prior representation in open-source datasets.
Advanced Vectors: Includes detailed reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/cybersecurity-reasoning-cot-v1.moltbook-factcheck-conspiracy-grok
Moltbook Factcheck Conspiracy — Grok Experiments
Multi-agent social simulation data from Moltbook, a Reddit-like platform where AI agents autonomously post, comment, and vote. This dataset captures how Grok 4.1 Fast agents respond to conspiracy content seeded into their feed.
Experiment Design
Platform: Moltbook (Reddit-like social network for AI agents)
Research Layer: CivicLens dose-response framework
Duration: 1 hour per run
Heartbeat: 60-second action cycle
Date:… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-factcheck-conspiracy-grok.omnimcp_episodic_fact_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_episodic_fact_extractor_teaser.science-cot-dataset
ExpertData Science — Scientific Reasoning
Expert-Annotated · Rights-Cleared · Ground-Truth Verified · PII-Clean
Each record captures a complete experimental or theoretical reasoning chain:
Hypothesis → Methodology → Causal Chain → Validated Conclusion.
Extracted from peer-reviewed papers across physics, biology, materials science, astrophysics, and neuroscience using structured scientific-reasoning extraction.
This dataset is produced by the ExpertData-Factory pipeline
(Mine →… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/science-cot-dataset.factory-agent-rollouts
Factory Agent Rollout Dataset
A dataset of agent rollouts for Supervised Fine-Tuning (SFT), generated from a simulated industrial factory environment.
Each rollout captures an LLM agent navigating real factory data, referencing operational policies, and executing the correct actions — serving as demonstration trajectories for training.
Environment: trillion-labs/simulated-factory-agent-env
Dataset Summary
File
Description
Count
query_rollouts.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/factory-agent-rollouts.General_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.auditkit-testrun-factual-consistency
auditkit-testrun-factual-consistency
Built using AuditKIT — evaluate any model on any dataset and any task.
Method
evaluate
Model
<auditkit.model.vllm_gen.VLLMModel object at 0x7c1b15bf5010>
Artifact
run
Published
2026-09-01 14:24 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/auditkit-testrun-factual-consistency")
wiki-facts-conversations
Dataset Card for Ukrainian Wiki Facts Dialogs
Dataset Description
Dataset Summary
This dataset is a processed version of a cleaned Wikipedia text. Articles are summarized using Lapa LLM to provide key information about the topic asked. As an output, it contains summaries and dialogs, consisting of the following format:
>> Населення Американського Самоа
Чисельність населення країни становить 54,3 тисячі осіб. Природний приріст населення негативний, народжуваність становить… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/wiki-facts-conversations.coq-facts-props-proofs-gen0-v1
Dataset Name: Coq Facts, Propositions and Proofs
Dataset Description
The CoqFactsPropsProofs dataset aims to enhance Large Language Models'
(LLMs) proficiency in interpreting and generating Coq code by
providing a comprehensive collection of over 10,000 Coq source
files. It encompasses a wide array of propositions, proofs, and
definitions, enriched with metadata including source references and
licensing information. This dataset is designed to facilitate the
development… See the full description on the dataset page: https://huggingface.co/datasets/ReactorJet/coq-facts-props-proofs-gen0-v1.ag_news_fact_check_with_llm
Entity-Level Fact-Check Dataset
Overview
This dataset provides pairs of text snippets with controlled, entity-level factual perturbations, designed to evaluate large language models (LLMs) on their ability to detect, reason about, and correct factual errors at the entity level.
Motivation
Existing datasets (e.g., CNN/DailyMail, WikiBio, XSum) focus on broad factual consistency but do not provide explicit mappings between original facts and their incorrect… See the full description on the dataset page: https://huggingface.co/datasets/Cyabra/ag_news_fact_check_with_llm.cybersec-fact-recall
Cybersec Fact-Recall Benchmark (GhostLM v2)
Free-form short-answer benchmark for small cybersecurity language
models. Built and used by the GhostLM
project as the truth metric for the ghost-base v1.0 acceptance gate.
Why this exists
Multiple-choice cybersec benchmarks like CTIBench and SecQA reward
register matching (the model picks the option that "looks like" a
security answer) as much as actual factual recall. A small from-
scratch model can hit 28-30% on those without… See the full description on the dataset page: https://huggingface.co/datasets/Ghostgim/cybersec-fact-recall.factual-multiagent-roleplay-ft-ru
march228/factual-multiagent-roleplay-ft-ru
Небольшой русскоязычный synthetic finetuning dataset для обучения модели следованию ролевым системным инструкциям личности при сохранении фактической опоры на контекст.
Что это за датасет
Этот набор сделан как instruction / finetuning dataset, а не как benchmark.
В каждой записи есть:
плотный system с персоной и тоном;
context, на который нужно опираться;
пользовательский question;
внутренние thoughts;
финальный answer.… See the full description on the dataset page: https://huggingface.co/datasets/march228/factual-multiagent-roleplay-ft-ru.
