datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical_meadow_pubmed_causal
Dataset Card for Pubmed Causal
Dataset Summary
This is the dataset used in the paper: Detecting Causal Language Use in Science Findings.
Citation Information
@inproceedings{yu-etal-2019-detecting,
title = "Detecting Causal Language Use in Science Findings",
author = "Yu, Bei and
Li, Yingya and
Wang, Jun",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_pubmed_causal.Synthetic-Causal-Reasoning-50k
🏭 Sovereign Synthetic Reasoning Dataset (400k)
"High-Quality Chain-of-Thought Data at Scale."
📊 Overview
This dataset contains 400,000 synthetic reasoning samples spanning 16 enterprise domains (Finance, Pharma, Legal, Cybersecurity, Supply Chain, etc.).
It was generated using the Sovereign Generator, which produced 1.6 million samples and applied a strict quality filter (Top 25%) to retain only the most logically consistent and complex chains.
Average Quality… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/Synthetic-Causal-Reasoning-50k.causalgymCausalGym is a benchmark for comparing the performance of causal interpretability methods
on a variety of simple linguistic tasks taken from the SyntaxGym evaluation set
(Gauthier et al., 2020, Hu et al., 2020)
and converted into a format suitable for interventional interpretability.
The dataset includes train/dev/test splits (exactly as used in the experiments in the paper).
The base/src columns are the prompts on which intervention is done. Each of these is a list of strings,
with each… See the full description on the dataset page: https://huggingface.co/datasets/aryaman/causalgym.causal-reasoning-benchmarks
Causal Reasoning Benchmarks
Datasets used in "On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning" (Deshmukh & Gupta, 2026).
Dataset Structure
train/transitivity_train.jsonl — 50,000 transitivity training examples
train/dsep_train.jsonl — 50,000 d-separation training examples
eval/length_eval.jsonl — 10,000 length generalization examples
eval/branching_eval.jsonl — 10,000 branching structure examples
eval/reversed_eval.jsonl — 10,000… See the full description on the dataset page: https://huggingface.co/datasets/ludwigw/causal-reasoning-benchmarks.repro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces
Agent traces
Agent sessions published from a Trackio Logbook.
vi-gym-causal-ascii
Vi-Gym Causal ASCII Trajectories
This dataset contains autoregressive trajectories of a Large Language Model (LLM) agent learning spatial reasoning and geometric drawing within a simulated Vi (Vim) editor environment.
Warning
This dataset is a direct derivation of the source material, it might therefore also contain content not suitable for all audiences. All authors of the original artwork have full ownership.
Dataset Structure
Each record is a discrete step… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/vi-gym-causal-ascii.stride-lds
STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations
Ground-truth Linear Datamodeling Score (LDS) targets for the four nanochat
pre-training models, plus the shared held-out test set.
Each lds_<tag>.jsonl was produced by: sampling a
pool of pre-training examples, drawing 256 random 30%-subsets, training a fresh nanochat
from scratch on each subset, and recording per-example held-out test losses. The
_meta header records the pool indices so scores defined over… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/stride-lds.GPT-4-Self-Instruct-GermanHere we share a German dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied.
We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly German. This dataset will be updated continuously.
CausalDriveBench
CausalDriveBench
A benchmark for causal reasoning in autonomous driving built on top of
nuScenes. Each sample bundles a curated
causal scene graph, three flavours of multiple-choice / open-ended QA
(active, dormant, distractor), and pointers to the raw nuScenes frames so the
benchmark stays compact and license-clean.
At a glance
Subset uploaded: nuscenes
Samples: 815 across 475 scenes
Tasks: active QA, dormant QA, distractor QA, causal scene graphs
Folder… See the full description on the dataset page: https://huggingface.co/datasets/causaldrivebench/CausalDriveBench.GPT-4-Self-Instruct-TurkishAs per the community's request, here we share a Turkish dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied.
We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly Turkish. This dataset will be… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/GPT-4-Self-Instruct-Turkish.GPT-4-Self-Instruct-JapaneseHere we share a Japanese dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied.
We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly Japanese. This dataset will be updated continuously.
CausalAirThe dataset is designed for fine-tuning a large language model for aviation accident analysis, including the reasoning process for identifying the causes of accidents. It consists of three fields: the narrative of the aviation accident, the reasoning process, and the accident cause. The dataset is divided into the SFT stage and the DPO stage.
GPT-4-Self-Instruct-GreekAs per the community's request, here we share a Greek dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied.
We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly Greek. This dataset will be… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/GPT-4-Self-Instruct-Greek.causal_kg_eval
Logic-Aware Causal Knowledge Graph Dataset
이 데이터셋은 비구조화된 텍스트에서 추출된 논리적 관계와 인과 경로의 타당성을 평가하기 위해 설계되었습니다. 단순한 유사도 기반 검색을 넘어, 구조화된 지식을 활용한 고난도 추론 성능 측정을 목적으로 합니다.
1. 개요 (Overview)
목적: 정보 간의 선후 관계, 인과성 및 다단계(Multi-hop) 연결성 검증
데이터 형식: JSON (head, relation, tail)
핵심 기능: Semantic Noise 필터링 및 논리적 추론 경로(Reasoning Path) 제공
2. 관계 스키마 (Relation Schema)
본 데이터셋은 정보 간의 연결 강도와 성격에 따라 다음 4가지 관계를 정의합니다.
Taxonomy (is_a): 상위 개념과 하위 개념 간의 계층적 분류
Causality (cause_of): 명확한 방향성을 가진… See the full description on the dataset page: https://huggingface.co/datasets/crjojo/causal_kg_eval.CausalLM__14B-details
Dataset Card for Evaluation run of CausalLM/14B
Dataset automatically created during the evaluation run of model CausalLM/14B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__14B-details.CausalLM__preview-1-hf-details
Dataset Card for Evaluation run of CausalLM/preview-1-hf
Dataset automatically created during the evaluation run of model CausalLM/preview-1-hf
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__preview-1-hf-details.Refined-Anime-Text
Refined Anime Text for Continual Pre-training of Language Models
This is a subset of our novel synthetic dataset of anime-themed text, containing over 1M entries, ~440M GPT-4/3.5 tokens. This dataset has never been publicly released before. We are releasing this subset due to the community's interest in anime culture, which is underrepresented in general-purpose datasets, and the low quality of raw text due to the prevalence of internet slang and irrelevant content, making it… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Refined-Anime-Text.Retrieval-SFT-Chat
Retrieval-Based Multi-Turn Chat SFT Synthetic Data
A year ago, we released CausalLM/Refined-Anime-Text, a thematic subset of a dataset generated using the then state-of-the-art LLMs. This dataset comprises 1 million entries synthesized through long-context models that rewrote multi-document web text inputs, intended for continued pre-training. We are pleased to note that this data has been employed in various training scenarios and in studies concerning data and internet culture.
In… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Retrieval-SFT-Chat.arayun_173-system-law-symbolic-causal-coherence
[DOI] https://doi.org/10.5281/zenodo.17186989
ARAYUN_173 – A System Law for Symbolic and Causal Coherence
Corresponding author: ARAYUN_173 (Independent Research)
E-mail: arayun173 [at] proton [dot] me
Website: arayun173.com
Date: September 2025
Audit Marker: SHA-256(ARAYUN_173|2025-09-04|Draft1)
Contact: arayun173 [at] proton [dot] me
ARAYUN_173 – A System Law for Symbolic and Causal Coherence
Abstract
ARAYUN_173 is not a concept but a system law. It establishes… See the full description on the dataset page: https://huggingface.co/datasets/ARAYUN173/arayun_173-system-law-symbolic-causal-coherence.causal_judgment
Causal Judgment: Causal Reasoning with Moral, Intentional, and Counterfactual Analysis
This task tests whether ultra-large language models are able to read a short story where multiple cause-and-effect events are introduced and answer causal questions such as "Did X cause Y?" in the same manner as humans would.
Authors: Allen Nie (anie@cs.stanford.edu), Tobias Gerstenberg (gerstenberg@stanford.edu)
Note: This repo is managed by the original author of this task.
Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/allenanie/causal_judgment.stride-preproc-climbmix
STRIDE: Preprocessed ClimbMix
Tokenized ClimbMix sequences used for STRIDE's pretraining-data attribution experiments. There is one training pool per released nanochat depth (d12, d16, d20, d24) and one held-out test set shared across depths. These are the corpora the released nanochat operators attribute over, so STRIDE's column indices line up with the lines of these files.
Files
File
Sequences
Size
Contents
climbmix_train_d12.jsonl
1,317,003
3.8 GB
training… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/stride-preproc-climbmix.CausalLM__34b-beta-details
Dataset Card for Evaluation run of CausalLM/34b-beta
Dataset automatically created during the evaluation run of model CausalLM/34b-beta
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__34b-beta-details.Causal-Conversationscausal-seal-pilot
Causal Seal — reference pilot
An illustrative, self-contained snapshot showing the full chain a third party can
follow without knowing the implementation that emitted it:
exact commit SHA
-> download / verify this snapshot
-> open the STS trace (trace.jsonl)
-> follow `causal_seal_ref` -> seal.json
-> install / run the published verifier
Files
file
role
manifest.json
entry point: what maps to what, sha256 per file. Git provides snapshot… See the full description on the dataset page: https://huggingface.co/datasets/causal-seal/causal-seal-pilot.DeepAutoAI__d2nwg_causal_gpt2_v1-details
Dataset Card for Evaluation run of DeepAutoAI/d2nwg_causal_gpt2_v1
Dataset automatically created during the evaluation run of model DeepAutoAI/d2nwg_causal_gpt2_v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DeepAutoAI__d2nwg_causal_gpt2_v1-details.NS-Causal-Inference-Syntheticdeu-medical_meadow_pubmed_causalDeepAutoAI__causal_gpt2-details
Dataset Card for Evaluation run of DeepAutoAI/causal_gpt2
Dataset automatically created during the evaluation run of model DeepAutoAI/causal_gpt2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DeepAutoAI__causal_gpt2-details.causal_datasets_mixedcausalab-graph-configs-review
CausaLab Causal Graph Configuration Dataset
Anonymous review release for NeurIPS 2026 Evaluations and Datasets.
This package contains the synthetic causal graph configuration suites used by the active experiments in the submitted CausaLab paper. Each JSONL file contains 50 graph configurations, one JSON object per line.
Contents
.
├── data/ # 19 JSONL graph-suite files, 950 total records
├── manifest.json # file purposes, record… See the full description on the dataset page: https://huggingface.co/datasets/CausaLab/causalab-graph-configs-review.
