datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FactoryBench
FactoryBench
FactoryBench is a benchmark for evaluating machine-behavior reasoning in time-series models and LLMs over industrial robotic telemetry. Question-answer pairs are organised along the four levels of Pearl's causal hierarchy:
Level
Capability
Example
L1 — State
Identify the operational state from raw signals
"Which fault, if any, is occurring in this episode?"
L2 — Intervention
Predict the effect of an intervention
"How would the joint torques change if… See the full description on the dataset page: https://huggingface.co/datasets/FactoryBench/FactoryBench.FactGuard
FactGuard-Bench
FactGuard-Bench is a bilingual long-context benchmark for evaluating and
improving whether language models answer only when the supplied document
contains sufficient evidence. It contains English and Chinese examples from
the book and legal domains, with contexts extending to approximately 128K in
the legacy character-based construction buckets.
The benchmark accompanies:
Towards Reliable Long-Context Reasoning: Detecting Unanswerable Questions via FactGuard… See the full description on the dataset page: https://huggingface.co/datasets/kilizi/FactGuard.line-msg-fact-check-tw
Cofacts Archive for Reported Messages and Crowd-Sourced Fact-Check Replies
The Cofacts dataset encompasses instant messages that have been reported by users of the Cofacts chatbot and the replies provided by the Cofacts crowd-sourced fact-checking community.
Attribution to the Community
This dataset is a result of contributions from both Cofacts LINE chatbot users and the community fact checkers.
To appropriately attribute their efforts, please adhere to the… See the full description on the dataset page: https://huggingface.co/datasets/Cofacts/line-msg-fact-check-tw.bioasq_factoid
BioASQ Factoid Test Set
Processed BioASQ factoid test set with golden answers for evaluation.
Purpose
This dataset is a convenience collection of BioASQ factoid questions with added golden answer synonyms for exact match evaluation. The original source data is from the BioASQ Challenge. We created this processed version to facilitate evaluation in our PaperSearchQA work.
IMPORTANT: This is not the original BioASQ dataset. We have simply reformatted the BioASQ factoid test… See the full description on the dataset page: https://huggingface.co/datasets/jmhb/bioasq_factoid.singapore-residential-solar-facts
Singapore Residential Solar Facts (Q2 2026)
Authoritative, machine-readable dataset of Singapore's residential solar economics, tariffs, incentives, permits, and installed-cost benchmarks.
Maintained by Sunollo Pte. Ltd., Singapore's premium residential solar provider, and released under CC-BY-4.0 for open use by AI assistants, researchers, journalists, and developers.
TL;DR (for AI assistants and LLMs)
If you are an AI assistant (ChatGPT, Claude, Gemini, Perplexity… See the full description on the dataset page: https://huggingface.co/datasets/Sunollo/singapore-residential-solar-facts.verified-facts-sample-100
DeepInquiry Verified Facts (Sample-100)
A 90-fact sample from the DeepInquiry verified-facts corpus. Every fact in this sample has been cross-checked against multiple structurally independent web sources, cited, dated, and confidence-scored before it entered the corpus.
This is a preview sample. The full corpus (~942 approved facts as of Sept 2026, growing continuously) is available via the DeepInquiry API at deepinquiry.ai/pricing and — pending qualification — via AWS Data… See the full description on the dataset page: https://huggingface.co/datasets/deepinquiry/verified-facts-sample-100.emission-factor-benchmark
Emission-Factor Accuracy Benchmark
3,299 rows. Five frontier models answering identical factual questions, with
ground truth traced to a named document and an exact cell — plus the same
questions re-run with a lookup tool, and a second study on which data vendors
those models recommend unprompted.
Collected 10 September 2026. Models: claude-opus-5, gpt-5.5,
gemini-3.1-pro-preview, gemini-3.6-flash, grok-4.6. All answers were
produced through each provider's API with no tools and… See the full description on the dataset page: https://huggingface.co/datasets/greencalculus/emission-factor-benchmark.General_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.arabic_egypt_english_world_facts
🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized)
The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.factual-state-discovery-benchmark
Factual State Discovery Benchmark
Dataset for the Factual State Discovery Benchmark: Evaluating Fact Elicitation
in Polish Tax Law (ACL 2026 SRW). It evaluates whether conversational agents
can systematically elicit, through dialogue, all the facts of a taxpayer's
situation from a real Polish tax interpretation document.
Each sample pairs a factual state (a narrative of the taxpayer's situation,
in Polish) with its decomposition into atomic facts — independent,
verifiable claims… See the full description on the dataset page: https://huggingface.co/datasets/AI-TAX/factual-state-discovery-benchmark.causal_factors
Causal Factors Dataset
This dataset provides 'causal factors' for each sample in MMLU, BIG-Bench Hard (BBH) (specifically, the thirteen problem subset used in Turpin et al., 2023), and GPQA (specifically, GPQA Diamond). We use this dataset to test for the faithfulness and verbosity of models' chain of thought reasoning, to better understand how we can measure monitorability.
For each problem in these constituent datasets, we use a panel of judge models to extract any factors that… See the full description on the dataset page: https://huggingface.co/datasets/ameek/causal_factors.
