datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 73 peer-reviewed research papers and three textbook-style collections (see CausalBenchmark.pdf). The benchmark contains 174 queries over 139 datasets. Each task is designed to evaluate both (i) identification, i.e., selecting an appropriate causal estimand and identification strategy given the study context, and… See the full description on the dataset page: https://huggingface.co/datasets/syrgkanislab/CausalReasoningBenchmark.causalverify-neurips2026
🎯 CausalVerify
An Execution-Grounded Benchmark for LLM Causal Inference Workflows
NeurIPS 2026 — Evaluations and Datasets Track · double-blind review · frozen at tag neurips2026-submission
💡 TL;DR
A benchmark of 259 published economics papers (Experiment A — real-paper text-agreement diagnostic) and 100 fixed-seed synthetic data-generating processes (Experiment B — execution-grounded coefficient recovery), evaluating 7 frontier LLMs. The central… See the full description on the dataset page: https://huggingface.co/datasets/causalverify/causalverify-neurips2026.CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Anonymized benchmark release. Author, affiliation, and prior-whitepaper material have been removed.
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 73 peer-reviewed research papers and three textbook-style collections. The benchmark contains 174 queries over 139 datasets. Each task is designed to evaluate both (i) identification, i.e., selecting an appropriate… See the full description on the dataset page: https://huggingface.co/datasets/anonsubmission16/CausalReasoningBenchmark.causal_factors
Causal Factors Dataset
This dataset provides 'causal factors' for each sample in MMLU, BIG-Bench Hard (BBH) (specifically, the thirteen problem subset used in Turpin et al., 2023), and GPQA (specifically, GPQA Diamond). We use this dataset to test for the faithfulness and verbosity of models' chain of thought reasoning, to better understand how we can measure monitorability.
For each problem in these constituent datasets, we use a panel of judge models to extract any factors that… See the full description on the dataset page: https://huggingface.co/datasets/ameek/causal_factors.stride-preproc-climbmix
STRIDE: Preprocessed ClimbMix
Tokenized ClimbMix sequences used for STRIDE's pretraining-data attribution experiments. There is one training pool per released nanochat depth (d12, d16, d20, d24) and one held-out test set shared across depths. These are the corpora the released nanochat operators attribute over, so STRIDE's column indices line up with the lines of these files.
Files
File
Sequences
Size
Contents
climbmix_train_d12.jsonl
1,317,003
3.8 GB
training… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/stride-preproc-climbmix.
