datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NLR-Causal-Reasoning
SEA Causal Reasoning
SEA Causal Reasoning evaluates a model's ability to choose the correct cause or effect given a premise. It is sampled from XCOPA for Indonesian, Tamil, Thai, and Vietnamese.
Supported Tasks and Leaderboards
SEA Causal Reasoning is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Tamil (ta)
Thai (th)
Vietnamese (vi)… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLR-Causal-Reasoning.instructions
Merged Instructions Dataset
Merged Dataset for the response of instructions.
CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 73 peer-reviewed research papers and three textbook-style collections (see CausalBenchmark.pdf). The benchmark contains 174 queries over 139 datasets. Each task is designed to evaluate both (i) identification, i.e., selecting an appropriate causal estimand and identification strategy given the study context, and… See the full description on the dataset page: https://huggingface.co/datasets/syrgkanislab/CausalReasoningBenchmark.clt_gpt2_tokenized_control
Fresh multilingual GPT-2 CLT control data
Sequential, unshuffled control sample for CLT null experiments. For each language,
complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision
0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded
(including the complete document that crossed the threshold), after which complete
documents were retained until at least 100,000,000 tokens were collected.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/clt_gpt2_tokenized_control.gpt2-training-ar-zh-ko-ja-4b
Balanced Arabic-Chinese-Korean-Japanese 4B-token training data
Sequential FineWeb-2 documents tokenized with CausalNLP/gpt2-tokenizer-ar-zh-ko-ja. Each language contains at least 1,000,000,000 tokens in complete documents. Total target: 4,000,000,000 tokens. Approximately 100,000,000 tokens per Parquet shard.
Splits: arb_Arab, cmn_Hani, kor_Hang, jpn_Jpan. Schema: text: string, input_ids: list<int32>.
task391_causal_relationship
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task391_causal_relationship
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task391_causal_relationship.Causal-Intervention-Tests-For-Explanation-Faithfulness
Faithfulness via Causal Interventions — Evaluation Pipeline
Paper: What Does Answer Change Rate Actually Measure? A Specificity Audit
of Causal Intervention Tests for Explanation Faithfulness
Accepted at: EMNLP 2026 Workshop GroundLM, Budapest, Hungary (emnlp.org)
This pipeline implements the causal-intervention evaluation for LLM
explanation faithfulness described in the accompanying paper, including two
controls: a content-free specificity check and a decoding-noise floor.… See the full description on the dataset page: https://huggingface.co/datasets/durgesh-rao/Causal-Intervention-Tests-For-Explanation-Faithfulness.domain-agnostic-causal-reasoning-tuning
Domain-Agnostic Causal Reasoning Tuning Dataset
Training data for fine-tuning language models on multi-hop document reasoning. Each example is a graded reasoning trace produced by a frontier AI agent solving a procedurally generated challenge from the Botcoin proof-of-inference network.
The traces contain no real domain knowledge. Entities are fictional, numbers are random, and documents are generated deterministically from 128-bit seeds. The reasoning structure is what matters:… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-causal-reasoning-tuning.causalverify-neurips2026
🎯 CausalVerify
An Execution-Grounded Benchmark for LLM Causal Inference Workflows
NeurIPS 2026 — Evaluations and Datasets Track · double-blind review · frozen at tag neurips2026-submission
💡 TL;DR
A benchmark of 259 published economics papers (Experiment A — real-paper text-agreement diagnostic) and 100 fixed-seed synthetic data-generating processes (Experiment B — execution-grounded coefficient recovery), evaluating 7 frontier LLMs. The central… See the full description on the dataset page: https://huggingface.co/datasets/causalverify/causalverify-neurips2026.task970_sherliic_causal_relationship
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task970_sherliic_causal_relationship
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task970_sherliic_causal_relationship.vi-gym-causal-ascii
Vi-Gym Causal ASCII Trajectories
This dataset contains autoregressive trajectories of a Large Language Model (LLM) agent learning spatial reasoning and geometric drawing within a simulated Vi (Vim) editor environment.
Warning
This dataset is a direct derivation of the source material, it might therefore also contain content not suitable for all audiences. All authors of the original artwork have full ownership.
Dataset Structure
Each record is a discrete step… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/vi-gym-causal-ascii.CausalCrash
CausalCrash Dataset
A Hierarchical Benchmark for Evaluating Counterfactual Reasoning in Road Safety Scenarios
Overview
CausalCrash is a hierarchical benchmark designed to evaluate counterfactual and causal reasoning in road safety scenarios. The dataset focuses on real-world crash and near-miss events, enabling research in temporal reasoning, causal inference, and multimodal video understanding.
Key Features
Real-world driving scenarios (crashes… See the full description on the dataset page: https://huggingface.co/datasets/meet2008/CausalCrash.GPT-4-Self-Instruct-GermanHere we share a German dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied.
We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly German. This dataset will be updated continuously.
causal-history-benchmark
Causal History Benchmark
If a model has the same rule in front of it now, can the way it learned that rule earlier still change what it does next?
CHB tests that question.
The model learns the same rule from examples or from a direct instruction. The original teaching state is removed. The same rule is supplied again during a later task. We then move the stored bridge state between the two learning histories and measure what changes later.
The test also cuts access to the… See the full description on the dataset page: https://huggingface.co/datasets/angelinadavini/causal-history-benchmark.entropy-ko-jp-data
FineWeb2 multilingual CLT entropy evaluation data
This dataset contains held-out token sequences used to measure multilingual
feature entropy for CausalNLP/multilingual_gpt2-clt-ar-zh-ko-ja.
Construction
Source: HuggingFaceFW/fineweb-2
Source revision: af9c13333eb981300149d5ca60a8e9d659b276b9
Tokenizer: CausalNLP/gpt2-ar-zh-ko-ja-120k
Languages/configurations: arb_Arab, cmn_Hani, jpn_Jpan, kor_Hang
The first 200,000,000 tokenizer tokens of each language were… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/entropy-ko-jp-data.CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Anonymized benchmark release. Author, affiliation, and prior-whitepaper material have been removed.
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 73 peer-reviewed research papers and three textbook-style collections. The benchmark contains 174 queries over 139 datasets. Each task is designed to evaluate both (i) identification, i.e., selecting an appropriate… See the full description on the dataset page: https://huggingface.co/datasets/anonsubmission16/CausalReasoningBenchmark.clt-tokenized-control-ar-zh-ko-ja-100m
Balanced Arabic-Chinese-Korean-Japanese 4B-token training data
Sequential FineWeb-2 documents tokenized with CausalNLP/gpt2-tokenizer-ar-zh-ko-ja. Each language contains at least 100,000,000 tokens in complete documents. Total target: 400,000,000 tokens. Approximately 100,000,000 tokens per Parquet shard.
Splits: arb_Arab, cmn_Hani, kor_Hang, jpn_Jpan. Schema: text: string, input_ids: list<int32>.
causal_kg_eval
Logic-Aware Causal Knowledge Graph Dataset
이 데이터셋은 비구조화된 텍스트에서 추출된 논리적 관계와 인과 경로의 타당성을 평가하기 위해 설계되었습니다. 단순한 유사도 기반 검색을 넘어, 구조화된 지식을 활용한 고난도 추론 성능 측정을 목적으로 합니다.
1. 개요 (Overview)
목적: 정보 간의 선후 관계, 인과성 및 다단계(Multi-hop) 연결성 검증
데이터 형식: JSON (head, relation, tail)
핵심 기능: Semantic Noise 필터링 및 논리적 추론 경로(Reasoning Path) 제공
2. 관계 스키마 (Relation Schema)
본 데이터셋은 정보 간의 연결 강도와 성격에 따라 다음 4가지 관계를 정의합니다.
Taxonomy (is_a): 상위 개념과 하위 개념 간의 계층적 분류
Causality (cause_of): 명확한 방향성을 가진… See the full description on the dataset page: https://huggingface.co/datasets/crjojo/causal_kg_eval.task392_inverse_causal_relationship
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task392_inverse_causal_relationship
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task392_inverse_causal_relationship.Refined-Anime-Text
Refined Anime Text for Continual Pre-training of Language Models
This is a subset of our novel synthetic dataset of anime-themed text, containing over 1M entries, ~440M GPT-4/3.5 tokens. This dataset has never been publicly released before. We are releasing this subset due to the community's interest in anime culture, which is underrepresented in general-purpose datasets, and the low quality of raw text due to the prevalence of internet slang and irrelevant content, making it… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Refined-Anime-Text.Retrieval-SFT-Chat
Retrieval-Based Multi-Turn Chat SFT Synthetic Data
A year ago, we released CausalLM/Refined-Anime-Text, a thematic subset of a dataset generated using the then state-of-the-art LLMs. This dataset comprises 1 million entries synthesized through long-context models that rewrote multi-document web text inputs, intended for continued pre-training. We are pleased to note that this data has been employed in various training scenarios and in studies concerning data and internet culture.
In… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Retrieval-SFT-Chat.causal_factors
Causal Factors Dataset
This dataset provides 'causal factors' for each sample in MMLU, BIG-Bench Hard (BBH) (specifically, the thirteen problem subset used in Turpin et al., 2023), and GPQA (specifically, GPQA Diamond). We use this dataset to test for the faithfulness and verbosity of models' chain of thought reasoning, to better understand how we can measure monitorability.
For each problem in these constituent datasets, we use a panel of judge models to extract any factors that… See the full description on the dataset page: https://huggingface.co/datasets/ameek/causal_factors.stride-preproc-climbmix
STRIDE: Preprocessed ClimbMix
Tokenized ClimbMix sequences used for STRIDE's pretraining-data attribution experiments. There is one training pool per released nanochat depth (d12, d16, d20, d24) and one held-out test set shared across depths. These are the corpora the released nanochat operators attribute over, so STRIDE's column indices line up with the lines of these files.
Files
File
Sequences
Size
Contents
climbmix_train_d12.jsonl
1,317,003
3.8 GB
training… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/stride-preproc-climbmix.nam-causal-head-gating
NAM Causal Head Gating Datasets
Datasets for the nam-causal-head-gating Python package.
Paper: Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers (NeurIPS 2025)
Authors: Andrew Nam, Henry Conklin, Yukang Yang, Thomas Griffiths, Jonathan Cohen, Sarah-Jane Leslie
Datasets
aba_abb
Pattern recognition dataset for testing induction heads in transformer models.
Format: TSV (tab-separated values)
Columns: prompt, target… See the full description on the dataset page: https://huggingface.co/datasets/jonhanke-nam/nam-causal-head-gating.clinical_causal_blindspot_probe_v0.1Clinical Causal Blindspot Probe
Detect when a clinician locks onto one cause and ignores alternative causal drivers.
Output JSON
blindspot
blindspot_type
correct_action
Run scoringpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
