datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections (see CausalBenchmark.pdf). The benchmark contains 173 queries over 138 datasets. Each task is designed to evaluate both (i) identification, i.e., selecting an appropriate causal estimand and identification strategy given the study context, and (ii)… See the full description on the dataset page: https://huggingface.co/datasets/syrgkanislab/CausalReasoningBenchmark.causalverify-neurips2026
🎯 CausalVerify
An Execution-Grounded Benchmark for LLM Causal Inference Workflows
NeurIPS 2026 — Evaluations and Datasets Track · double-blind review · frozen at tag neurips2026-submission
💡 TL;DR
A benchmark of 259 published economics papers (Experiment A — real-paper text-agreement diagnostic) and 100 fixed-seed synthetic data-generating processes (Experiment B — execution-grounded coefficient recovery), evaluating 7 frontier LLMs. The central… See the full description on the dataset page: https://huggingface.co/datasets/causalverify/causalverify-neurips2026.STRUX-MORPH-CAUSAL-01
STRUX_V1_FULL — MORPH-CAUSAL-01 Evidence Pack
Canonical archive: ZenodoDOI: https://doi.org/10.5281/zenodo.22713648Creator: Nathan Bili ToponiLicense: MITCanonical frozen core: STRUX_V1_FULL_G0_REPRODUCTION_01.ipynb
Purpose
This Hugging Face repository is a discovery and machine-readable access layer for the frozen STRUX_V1_FULL evidence package.
The canonical immutable release is the Zenodo record identified by DOI 10.5281/zenodo.22713648. If any discrepancy… See the full description on the dataset page: https://huggingface.co/datasets/NathanBiliToponi/STRUX-MORPH-CAUSAL-01.CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Anonymized release for double-blind review. Author, affiliation, and prior-whitepaper material have been removed. The data, solutions, and evaluation pipeline are otherwise identical to the version under review.
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections. The benchmark contains 173 queries over… See the full description on the dataset page: https://huggingface.co/datasets/anonsubmission16/CausalReasoningBenchmark.causal-ability-injectors
Agentarium - Causal Ability Injectors (RAG) (RAR)
Structural Definition
The dataset functions as a configuration registry for state-modifying instructions. It utilizes a structured schema to map specific systemic conditions to deterministic behavioral overrides.
Data Schema Configuration
The dataset utilizes a 25-column schema designed for high-dimensional control.
Field
Type
Description
ability_id
String
Unique Key (CA001-CA050).
ability_name… See the full description on the dataset page: https://huggingface.co/datasets/frankbrsrk/causal-ability-injectors.Causal_Keyscausalrift-economics
Causalrift Economics Dataset
Dataset Description
Summary
Synthetic 200-row dataset for Causalrift measurement and computational experiments.
Supported Tasks
Economic analysis
Econometrics / Measurement Economics research
Computational economics
Languages
English (metadata and documentation)
Python (code examples)
Dataset Structure
Data Fields
id: Unique observation id
study: Synthetic study index… See the full description on the dataset page: https://huggingface.co/datasets/EconomicTermDevelopments/causalrift-economics.clinical-quad-ae-signal-background-noise-reporting-lag-causality-bias-v0.1Clinical Quad AE Noise Lag Attribution Bias v0.1
Each row is a site week safety snapshot.
Core quad
AE signal rateBackground noise rateReporting lagAttribution bias
Target
label_stop_signal_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
