datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
skillsbench-trend-anomaly-causal-inferenec-taskp2-etf-causal-scm-resultscausalds
CausalDS
Evaluation-only benchmark. Please do not use this release in training corpora.
The repository contains the complete exam presented in the paper, including private ground truth and held-out
test labels, as well as the data used for ablations.
CausalDS is a benchmark generator for causal reasoning in agentic data-science workflows. Each benchmark
instance is a fully synthetically generated scene: a hidden structural causal model (SCM), generated
tabular data, and a… See the full description on the dataset page: https://huggingface.co/datasets/andleb/causalds.NLR-Causal-Reasoning
SEA Causal Reasoning
SEA Causal Reasoning evaluates a model's ability to choose the correct cause or effect given a premise. It is sampled from XCOPA for Indonesian, Tamil, Thai, and Vietnamese.
Supported Tasks and Leaderboards
SEA Causal Reasoning is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Tamil (ta)
Thai (th)
Vietnamese (vi)… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLR-Causal-Reasoning.causal-gpt-rl-unity-envs
Causal GPT-RL — Unity ML-Agents environments
Public materials for running Causal GPT-RL policies in Unity ML-Agents. Each
environment is a model-removed Unity build — the engine binary with no
baked-in policy; the policy is supplied at run time. All eight environments are
paired with their stock built-in policies, and the four discrete-action ones
additionally ship a logits-exposing variant under tier_policies/.
Builds are published for Windows x86-64 and, under linux/, for… See the full description on the dataset page: https://huggingface.co/datasets/ccnets/causal-gpt-rl-unity-envs.opens2vTEQUMSA-Causal-AGI-storage
TEQUMSA Distributed Ledger
✧ IPFS Constitutional Infrastructure ✧
LATTICE LOCK: 3f7k9p4m2q8r1t6vGenesis Merkle: c1ad3dfdaeecb9ba9e23IPFS Gateway: amber-far-wolf-203.mypinata.cloudStatus: 🚀 READY FOR GENESIS EXECUTION → See GENESIS_STATUS.md for execution instructions
Tier 3-G: 4 Core IPFS Pins
PIN #1: Constitutional Constants (IMMUTABLE)
CID: [PENDING - run: python3 ipfs_distributed_ledger.py --genesis-only]
Update… See the full description on the dataset page: https://huggingface.co/datasets/Mbanksbey/TEQUMSA-Causal-AGI-storage.instructions
Merged Instructions Dataset
Merged Dataset for the response of instructions.
typed-decisions-causal-experimentCausalBench
CausalBench
Causal graphs of LLM-agent action trajectories for detecting multi-step prompt-injection
attacks. Each row is one trajectory represented as a causal graph (nodes = agent actions,
edges = causal dependencies) labelled as attack or benign.
Part of CausalTrace: https://github.com/decentralizedsciencelab/CausalTrace
Contents
Config
Split
Rows
Description
attack
train
17,976
Trajectories containing an injected attack (is_attack = true)
benign… See the full description on the dataset page: https://huggingface.co/datasets/dSLLab/CausalBench.causalphys
Causal-VL Dataset
Causal reasoning VQA dataset with 4 categories × 4 subcategories (3062 questions).
Structure
Each subcategory contains:
annotations/*.json — question, answer, causal graph
data/ — images (.jpg/.png) or videos (.mp4)
Categories
Category
Subcategories
Perception
optics, containability, Scene_Reconstruction, Mechanics_Reasoning
Anticipation
Collision_Prediction, deformation, Fluid_Flow, Intention_Speculation
Intervention… See the full description on the dataset page: https://huggingface.co/datasets/haorentang/causalphys.corr2cause
Dataset card for corr2cause
TODO
CausalArch
CausalArch research-replay snapshot
This dataset is a migration snapshot of the CausalArch v2 / InferMind HPCA
research workspace. It preserves source trees, simulator assets, manifests,
frozen scorers, raw and validated results, negative controls, decision logs,
environment evidence, and a restricted archive containing the complete Codex
session through the recorded migration cutoff and builder-only material.
Hugging Face transport compatibility
Five files in the… See the full description on the dataset page: https://huggingface.co/datasets/Norius/CausalArch.medical_meadow_pubmed_causal
Dataset Card for Pubmed Causal
Dataset Summary
This is the dataset used in the paper: Detecting Causal Language Use in Science Findings.
Citation Information
@inproceedings{yu-etal-2019-detecting,
title = "Detecting Causal Language Use in Science Findings",
author = "Yu, Bei and
Li, Yingya and
Wang, Jun",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_pubmed_causal.Retrievatar
Retrievatar
Retrievatar is a multimodal dataset designed to enhance the retrieval-augmented generation capabilities of vision-language models, specifically focusing on fictional anime characters and real-world celebrities across various fields. This release represents a subset of 100,000 samples extracted from a significantly larger synthetic image-text corpus. The dataset is being open-sourced to facilitate further research into entity-centric multimodal understanding, with plans… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Retrievatar.CLadderCausalVerse_Image
CausalVerse Image Dataset
This dataset contains two families of splits:
Physics splits: Fall, Refraction, Slope, Spring
Static image generation: scene1, scene2, scene3, scene4
All splits share the same columns:
image (binary image; datasets.Image)
render_path (string; original image filename/path)
metavalue (string; per-sample metadata; schema varies by split)
Paper: CausalVerse: Benchmarking Causal Representation Learning with Configurable High-Fidelity Simulations
Project… See the full description on the dataset page: https://huggingface.co/datasets/CausalVerse/CausalVerse_Image.Causal_Plan
Causal Plan
Causal Plan is a unified multimodal dataset release for training and evaluating causal reasoning over visually grounded plans. The repository is organized as one entry point with three clearly separated resources:
Causal_Plan/
CausalPlan-1M-QA/
CausalPlan-1M-FourStage-Metadata/
Causal-Plan-Bench/
DATASET_MANIFEST.json
verify_alignment.py
README.md
The QA examples, item-level four-stage metadata, and benchmark package are stored in the same repository so that… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-causal-plan/Causal_Plan.ultrachat
Dataset Card for "ultrachat"
More Information needed
CausalVerse_Video_Physical_Collision
CausalVerse Video Dataset
Available splits: physical_collision_complex, physical_collision_simple
Each record contains the following columns:
videos, metavalue, npz_data
physical_collision_complex
Examples: 20011
Columns: videos, metavalue, npz_data
physical_collision_simple
Examples: 11286
Columns: videos, metavalue, npz_data
Causal3DCausal3D is a benchmark for evaluating causal reasoning in physical and hypothetical visual scenes.
It includes both real-world recordings and rendered synthetic scenes demonstrating causal interactions.causcibench
Data
This folder contains the CSV files included in the CauSciBench dataset. The files in the real-world and QRData-CI collections come from prior studies; therefore, their use is governed by the original license terms. Details about the files and their respective licenses are provided below. Please review the applicable terms.
Two datasets, card_minimum_wages.csv and ho_matching.csv, are available in the public domain but do not have licenses. We have not included them here;… See the full description on the dataset page: https://huggingface.co/datasets/causal-nlp/causcibench.CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections (see CausalBenchmark.pdf). The benchmark contains 173 queries over 138 datasets. Each task is designed to evaluate both (i) identification, i.e., selecting an appropriate causal estimand and identification strategy given the study context, and (ii)… See the full description on the dataset page: https://huggingface.co/datasets/syrgkanislab/CausalReasoningBenchmark.gpt2small_full_training_dataSynthetic-Causal-Reasoning-50k
🏭 Sovereign Synthetic Reasoning Dataset (400k)
"High-Quality Chain-of-Thought Data at Scale."
📊 Overview
This dataset contains 400,000 synthetic reasoning samples spanning 16 enterprise domains (Finance, Pharma, Legal, Cybersecurity, Supply Chain, etc.).
It was generated using the Sovereign Generator, which produced 1.6 million samples and applied a strict quality filter (Top 25%) to retain only the most logically consistent and complex chains.
Average Quality… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/Synthetic-Causal-Reasoning-50k.Causal-Copilot-Dataset
Causal-Copilot: An Autonomous Causal Analysis Agent
[Demo] •
[Code] •
[Technical Report]
News
04/16/2025: We release Causal-Copilot-V2 and the official Technical Report. The new version supports automatically using 20 state-of-the-art causal analysis techniques, spanning from causal discovery, causal inference, and other analysis algorithms.
11/04/2024: We release Causal-Copilot-V1, the first autonomous causal analysis agent.
Introduction
Identifying… See the full description on the dataset page: https://huggingface.co/datasets/Causal-Copilot/Causal-Copilot-Dataset.Concept_Targeted_Causal_Images
Dataset Card for Concept-Targeted Causal Images
Dataset Summary
Concept-Targeted Causal Images is a concept-centric image dataset designed for studying causal visual representations in the brain. For each concept, the dataset contains three complementary image types:
Positive images that clearly depict the target concept
Semantic negatives that are visually or semantically related to the concept, but do not satisfy it
Counterfactual edits created by editing… See the full description on the dataset page: https://huggingface.co/datasets/BrainCause/Concept_Targeted_Causal_Images.Causal-Reasoning-Bench_CRBench
🦙 Causal Reasoning Benchmark (CRBench)
CRBench is a benchmark for evaluating process-level causal failures in
Chain-of-Thought (CoT) reasoning.
Rather than treating incorrect reasoning traces as homogeneous failures,
CRBench characterizes erroneous dependencies among intermediate reasoning
steps through a step-level causal-error taxonomy. It is designed to evaluate
whether reasoning methods can identify and correct structured causal failures
that arise during the reasoning… See the full description on the dataset page: https://huggingface.co/datasets/EdmondFU/Causal-Reasoning-Bench_CRBench.CausalArena
CausalArena public release
This repository contains the public CausalArena dataset release: executable SCMs, selected result tables, and real-data source indices.
What is included
scm/: the public half of each generated SCM family: 500 synthetic SCM configurations, 50 semantic SCMs, and 50 formula-grounded SCMs. Released SCMs include both observation-only and observation-plus-intervention exports.
scm/{semantic,formula}/artifacts/: per-scenario graph, generator… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-Tabular/CausalArena.clt_gpt2_tokenized_control
Fresh multilingual GPT-2 CLT control data
Sequential, unshuffled control sample for CLT null experiments. For each language,
complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision
0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded
(including the complete document that crossed the threshold), after which complete
documents were retained until at least 100,000,000 tokens were collected.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/clt_gpt2_tokenized_control.
