datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
causalds
CausalDS
Evaluation-only benchmark. Please do not use this release in training corpora.
The repository contains the complete exam presented in the paper, including private ground truth and held-out
test labels, as well as the data used for ablations.
CausalDS is a benchmark generator for causal reasoning in agentic data-science workflows. Each benchmark
instance is a fully synthetically generated scene: a hidden structural causal model (SCM), generated
tabular data, and a… See the full description on the dataset page: https://huggingface.co/datasets/andleb/causalds.CausalBench
CausalBench
Causal graphs of LLM-agent action trajectories for detecting multi-step prompt-injection
attacks. Each row is one trajectory represented as a causal graph (nodes = agent actions,
edges = causal dependencies) labelled as attack or benign.
Part of CausalTrace: https://github.com/decentralizedsciencelab/CausalTrace
Contents
Config
Split
Rows
Description
attack
train
17,976
Trajectories containing an injected attack (is_attack = true)
benign… See the full description on the dataset page: https://huggingface.co/datasets/dSLLab/CausalBench.corr2cause
Dataset card for corr2cause
TODO
CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections (see CausalBenchmark.pdf). The benchmark contains 173 queries over 138 datasets. Each task is designed to evaluate both (i) identification, i.e., selecting an appropriate causal estimand and identification strategy given the study context, and (ii)… See the full description on the dataset page: https://huggingface.co/datasets/syrgkanislab/CausalReasoningBenchmark.Causal2Needles
Causal2Needles (NeurIPS D&B Track 2025)
Overview
Project
Paper
Code
Causal2Needles is a benchmark dataset and evaluation toolkit designed to assess the capabilities of both proprietary and open-source multimodal large language models in long-video understanding. Our dataset features a large number of "2-needle" questions, where the model must locate and reason over two distinct pieces of information from the video. An illustrative example is shown below:
More background… See the full description on the dataset page: https://huggingface.co/datasets/causal2needles/Causal2Needles.details_CausalLM__34b-beta
Dataset Card for Evaluation run of CausalLM/34b-beta
Dataset automatically created during the evaluation run of model CausalLM/34b-beta.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_CausalLM__34b-beta.causalverify-neurips2026
🎯 CausalVerify
An Execution-Grounded Benchmark for LLM Causal Inference Workflows
NeurIPS 2026 — Evaluations and Datasets Track · double-blind review · frozen at tag neurips2026-submission
💡 TL;DR
A benchmark of 259 published economics papers (Experiment A — real-paper text-agreement diagnostic) and 100 fixed-seed synthetic data-generating processes (Experiment B — execution-grounded coefficient recovery), evaluating 7 frontier LLMs. The central… See the full description on the dataset page: https://huggingface.co/datasets/causalverify/causalverify-neurips2026.repro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces
Agent traces
Agent sessions published from a Trackio Logbook.
stride-lds
STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations
Ground-truth Linear Datamodeling Score (LDS) targets for the four nanochat
pre-training models, plus the shared held-out test set.
Each lds_<tag>.jsonl was produced by: sampling a
pool of pre-training examples, drawing 256 random 30%-subsets, training a fresh nanochat
from scratch on each subset, and recording per-example held-out test losses. The
_meta header records the pool indices so scores defined over… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/stride-lds.ultrafeedback_improve-degrade_qrandomized-neutrals-100k_ca-neutrals-on-causals-100kaurelia-federation-causal
Aurelia — Federation Causal Graph (Events + Edges)
Cross-world and federation-level causal events (events.parquet) plus the causal edges (edges.parquet) that connect them. The federation graph captures trade shocks, migrations, cultural diffusion, and diplomatic relations across the five Aurelian worlds. Edges are filtered to the federation events in this run.
Provenance
This dataset was generated by Aurelia,
a multi-agent civilization simulation operated by Ousia… See the full description on the dataset page: https://huggingface.co/datasets/OusiaResearch/aurelia-federation-causal.synthetic_causal_pairsCausal pairs generated with chatGPT. Training set.
STRUX-MORPH-CAUSAL-01
STRUX_V1_FULL — MORPH-CAUSAL-01 Evidence Pack
Canonical archive: ZenodoDOI: https://doi.org/10.5281/zenodo.22713648Creator: Nathan Bili ToponiLicense: MITCanonical frozen core: STRUX_V1_FULL_G0_REPRODUCTION_01.ipynb
Purpose
This Hugging Face repository is a discovery and machine-readable access layer for the frozen STRUX_V1_FULL evidence package.
The canonical immutable release is the Zenodo record identified by DOI 10.5281/zenodo.22713648. If any discrepancy… See the full description on the dataset page: https://huggingface.co/datasets/NathanBiliToponi/STRUX-MORPH-CAUSAL-01.ultrafeedback_improve-degrade_qrandomized-neutrals-60k_ca-neutrals-on-causals-60kCausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Anonymized release for double-blind review. Author, affiliation, and prior-whitepaper material have been removed. The data, solutions, and evaluation pipeline are otherwise identical to the version under review.
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections. The benchmark contains 173 queries over… See the full description on the dataset page: https://huggingface.co/datasets/anonsubmission16/CausalReasoningBenchmark.stride-preproc-climbmix
STRIDE: Preprocessed ClimbMix
Tokenized ClimbMix sequences used for STRIDE's pretraining-data attribution experiments. There is one training pool per released nanochat depth (d12, d16, d20, d24) and one held-out test set shared across depths. These are the corpora the released nanochat operators attribute over, so STRIDE's column indices line up with the lines of these files.
Files
File
Sequences
Size
Contents
climbmix_train_d12.jsonl
1,317,003
3.8 GB
training… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/stride-preproc-climbmix.streaming-gebd-causal
Audited Kinetics-GEBD Causal Metadata
This metadata-only release converts the publicly released Kinetics-GEBD
annotations into an auditable 24 FPS causal training representation. It does
not redistribute Kinetics or YouTube video bytes.
Splits
Hub split
Official source file
Records
Meaning
train
k400_train_raw_annotation.pkl
18,808
Public GEBD training annotations
validation
k400_val_raw_annotation.pkl
18,815
Public Kinetics-GEBD validation… See the full description on the dataset page: https://huggingface.co/datasets/kfkas/streaming-gebd-causal.pick_and_place_ball_causal_trainThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.front": {
"dtype": "video",
"shape": [
240,
320,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/ForsythYa/pick_and_place_ball_causal_train.causalrec-bench
CausalRec-Bench
A fully synthetic, multi-domain (e-commerce and streaming) recommendation interaction
dataset generated from a documented, exactly reproducible structural model. Every click
carries literal counterfactual confounder attribution — derived by reusing the same
random draw that generated the observed outcome to test which confounders were causally
necessary — rather than a post-hoc magnitude heuristic.
Paper: CausalRec-Bench: A Synthetic Benchmark with Counterfactual… See the full description on the dataset page: https://huggingface.co/datasets/alihassan1437/causalrec-bench.CausalLM__14B-details
Dataset Card for Evaluation run of CausalLM/14B
Dataset automatically created during the evaluation run of model CausalLM/14B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__14B-details.swedish-causality-binary
Swedish Causality Binary Classification Dataset
Binary causality detection dataset for Swedish text, extracted from Swedish Government Official Reports (SOU-corpus).
Dataset Description
This dataset contains Swedish sentences annotated for the presence of causal relations. Each example includes:
theme: The thematic category (e.g., "skog, växthuseffekt/klimat")
left_context: Preceding context sentences
target_sentence: The sentence to classify
right_context: Following… See the full description on the dataset page: https://huggingface.co/datasets/UppsalaNLP/swedish-causality-binary.pick_and_place_ball_causal_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.front": {
"dtype": "video",
"shape": [
240,
320,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/ForsythYa/pick_and_place_ball_causal_test.CausalLM__preview-1-hf-details
Dataset Card for Evaluation run of CausalLM/preview-1-hf
Dataset automatically created during the evaluation run of model CausalLM/preview-1-hf
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__preview-1-hf-details.super_poulain_draft_contactsheet_causalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 50,
"total_frames": 32650,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/super_poulain_draft_contactsheet_causal.ultrafeedback_60658_preference_dataset_original_neutrals_filtered_improve-degrade_filtered0p2chess-causal-formattedultrafeedback_improve-degrade_ca-neutrals-on-causals-100kultrafeedback_improve-degrade_qrandomized-neutrals-30k_ca-neutrals-on-causals-30kgaleras-causal4se-3k-levenshteindms-causal-pathways-data
