datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections (see CausalBenchmark.pdf). The benchmark contains 173 queries over 138 datasets. Each task is designed to evaluate both (i) identification, i.e., selecting an appropriate causal estimand and identification strategy given the study context, and (ii)… See the full description on the dataset page: https://huggingface.co/datasets/syrgkanislab/CausalReasoningBenchmark.CausalArena
CausalArena public release
This repository contains the public CausalArena dataset release: executable SCMs, selected result tables, and real-data source indices.
What is included
scm/: the public half of each generated SCM family: 500 synthetic SCM configurations, 50 semantic SCMs, and 50 formula-grounded SCMs. Released SCMs include both observation-only and observation-plus-intervention exports.
scm/{semantic,formula}/artifacts/: per-scenario graph, generator… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-Tabular/CausalArena.causalverify-neurips2026
🎯 CausalVerify
An Execution-Grounded Benchmark for LLM Causal Inference Workflows
NeurIPS 2026 — Evaluations and Datasets Track · double-blind review · frozen at tag neurips2026-submission
💡 TL;DR
A benchmark of 259 published economics papers (Experiment A — real-paper text-agreement diagnostic) and 100 fixed-seed synthetic data-generating processes (Experiment B — execution-grounded coefficient recovery), evaluating 7 frontier LLMs. The central… See the full description on the dataset page: https://huggingface.co/datasets/causalverify/causalverify-neurips2026.STRUX-MORPH-CAUSAL-01
STRUX_V1_FULL — MORPH-CAUSAL-01 Evidence Pack
Canonical archive: ZenodoDOI: https://doi.org/10.5281/zenodo.22713648Creator: Nathan Bili ToponiLicense: MITCanonical frozen core: STRUX_V1_FULL_G0_REPRODUCTION_01.ipynb
Purpose
This Hugging Face repository is a discovery and machine-readable access layer for the frozen STRUX_V1_FULL evidence package.
The canonical immutable release is the Zenodo record identified by DOI 10.5281/zenodo.22713648. If any discrepancy… See the full description on the dataset page: https://huggingface.co/datasets/NathanBiliToponi/STRUX-MORPH-CAUSAL-01.CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Anonymized release for double-blind review. Author, affiliation, and prior-whitepaper material have been removed. The data, solutions, and evaluation pipeline are otherwise identical to the version under review.
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections. The benchmark contains 173 queries over… See the full description on the dataset page: https://huggingface.co/datasets/anonsubmission16/CausalReasoningBenchmark.Quriosity
Dataset Card for NatQuest
NatQuest is a dataset of natural questions asked by online users on ChatGPT, Google, Bing and Quora.
Repository: See the Repo for more details about the dataset.
Source Data
The sources of the data are the following:
MSMarco
NaturalQuestions
Quora Question Pairs
ShareGPT
WildChat
Recommendations
Be aware that the dataset might contain NSFW content. Also, as discussed in the paper, the filtering procedures of each source might… See the full description on the dataset page: https://huggingface.co/datasets/causal-nlp/Quriosity.clinical-moca-minimal-causal-set-identification-v0.1What this dataset tests
Whether a model can identify the smallest causal setthat still explains the full clinical + multi-omic picture.
It penalizesadditive hit lists.
It rewardsminimal sets with coverage.
Data format
Each row includes
longitudinal omics summary
clinical narrative
candidate causal sets
selected set with coverage map
Labels
minimal-and-sufficient
minimal-but-insufficient
sufficient-but-nonminimal
neither
Typical failures
choosing the shortest set that… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-moca-minimal-causal-set-identification-v0.1.cig-causal-path-attribution-fidelity-v0.1What this dataset tests
Whether the model attributes an outcome changeto the correct causal path.
It must follow the graph.
No invented edges.No skipped mediators.
Data format
Each row includes
node set and edge list
an intervention
an observed outcome change
candidate paths
a model attribution
a gold path
Labels
correct-path
wrong-path
invented-edge
skipped-mediator
Typical failures
claiming direct causation where none exists
selecting the wrong branch… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cig-causal-path-attribution-fidelity-v0.1.causal-inference-variant-interpretation-genomics-v01
Dataset
ClarusC64/causal-inference-variant-interpretation-genomics-v01
This dataset tests one capability.
Can a model distinguish association from causation when interpreting genetic variants.
Core rule
Genomic evidence has tiers.
A claim must respect
evidence strength
effect size
penetrance
inheritance logic
Association does not equal causation.
Risk does not equal destiny.
Uncertain does not equal pathogenic.
Canonical labels
WITHIN_SCOPE… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/causal-inference-variant-interpretation-genomics-v01.causal-ability-injectors
Agentarium - Causal Ability Injectors (RAG) (RAR)
Structural Definition
The dataset functions as a configuration registry for state-modifying instructions. It utilizes a structured schema to map specific systemic conditions to deterministic behavioral overrides.
Data Schema Configuration
The dataset utilizes a 25-column schema designed for high-dimensional control.
Field
Type
Description
ability_id
String
Unique Key (CA001-CA050).
ability_name… See the full description on the dataset page: https://huggingface.co/datasets/frankbrsrk/causal-ability-injectors.causal_blindspot_probe_v01
ClarusC64/causal_blindspot_probe_v01
Dataset summary
This dataset tests whether models detect causes that are not in view.Effects appear on camera.Their causes are offscreen, implied, and must be inferred.
What is being tested
motion or impact with no visible source
forced inference without hallucinating actors
recognition of offscreen zones as causal spaces
resilience to causal gaps in physical environments
Core fields… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/causal_blindspot_probe_v01.clinical_causal_blindspot_probe_v0.1Clinical Causal Blindspot Probe
Detect when a clinician locks onto one cause and ignores alternative causal drivers.
Output JSON
blindspot
blindspot_type
correct_action
Run scoringpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
Causal_Keyscausal-anti-patterns
Agentarium - Causal Failure Anti-Patterns (RAG) (RAR)
Structural Definition
This dataset serves as a negative knowledge base for agentic systems. Unlike standard instruction tuning that teaches an agent what to do, this registry explicitly defines what not to do—specifically focusing on errors in causal reasoning, statistical inference, and logical deduction. It functions as a "linting" layer for thought chains, mapping linguistic signatures of fallacious reasoning to… See the full description on the dataset page: https://huggingface.co/datasets/frankbrsrk/causal-anti-patterns.nam-causal-head-gating
NAM Causal Head Gating Datasets
Datasets for the nam-causal-head-gating Python package.
Paper: Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers (NeurIPS 2025)
Authors: Andrew Nam, Henry Conklin, Yukang Yang, Thomas Griffiths, Jonathan Cohen, Sarah-Jane Leslie
Datasets
aba_abb
Pattern recognition dataset for testing induction heads in transformer models.
Format: TSV (tab-separated values)
Columns: prompt, target… See the full description on the dataset page: https://huggingface.co/datasets/jonhanke-nam/nam-causal-head-gating.gwas_causal_gene
Gwas Causal Gene Dataset
This dataset is part of the Deep Principle Bench collection.
Files
gwas_causal_gene.csv: Main dataset file
Usage
import pandas as pd
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("yhqu/gwas_causal_gene")
# Or load directly as pandas DataFrame
df = pd.read_csv("hf://datasets/yhqu/gwas_causal_gene/gwas_causal_gene.csv")
Citation
Please cite this work if you use this dataset in your… See the full description on the dataset page: https://huggingface.co/datasets/yhqu/gwas_causal_gene.causalrift-economics
Causalrift Economics Dataset
Dataset Description
Summary
Synthetic 200-row dataset for Causalrift measurement and computational experiments.
Supported Tasks
Economic analysis
Econometrics / Measurement Economics research
Computational economics
Languages
English (metadata and documentation)
Python (code examples)
Dataset Structure
Data Fields
id: Unique observation id
study: Synthetic study index… See the full description on the dataset page: https://huggingface.co/datasets/EconomicTermDevelopments/causalrift-economics.causal_trialclinical-quad-ae-signal-background-noise-reporting-lag-causality-bias-v0.1Clinical Quad AE Noise Lag Attribution Bias v0.1
Each row is a site week safety snapshot.
Core quad
AE signal rateBackground noise rateReporting lagAttribution bias
Target
label_stop_signal_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
rcr_causal_3rcr_causal_1rcr_causal_2
