datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NLR-Causal-Reasoning
SEA Causal Reasoning
SEA Causal Reasoning evaluates a model's ability to choose the correct cause or effect given a premise. It is sampled from XCOPA for Indonesian, Tamil, Thai, and Vietnamese.
Supported Tasks and Leaderboards
SEA Causal Reasoning is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Tamil (ta)
Thai (th)
Vietnamese (vi)… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLR-Causal-Reasoning.instructions
Merged Instructions Dataset
Merged Dataset for the response of instructions.
CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections (see CausalBenchmark.pdf). The benchmark contains 173 queries over 138 datasets. Each task is designed to evaluate both (i) identification, i.e., selecting an appropriate causal estimand and identification strategy given the study context, and (ii)… See the full description on the dataset page: https://huggingface.co/datasets/syrgkanislab/CausalReasoningBenchmark.clt_gpt2_tokenized_control
Fresh multilingual GPT-2 CLT control data
Sequential, unshuffled control sample for CLT null experiments. For each language,
complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision
0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded
(including the complete document that crossed the threshold), after which complete
documents were retained until at least 100,000,000 tokens were collected.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/clt_gpt2_tokenized_control.gpt2-training-ar-zh-ko-ja-4b
Balanced Arabic-Chinese-Korean-Japanese 4B-token training data
Sequential FineWeb-2 documents tokenized with CausalNLP/gpt2-tokenizer-ar-zh-ko-ja. Each language contains at least 1,000,000,000 tokens in complete documents. Total target: 4,000,000,000 tokens. Approximately 100,000,000 tokens per Parquet shard.
Splits: arb_Arab, cmn_Hani, kor_Hang, jpn_Jpan. Schema: text: string, input_ids: list<int32>.
task391_causal_relationship
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task391_causal_relationship
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task391_causal_relationship.domain-agnostic-causal-reasoning-tuning
Domain-Agnostic Causal Reasoning Tuning Dataset
Training data for fine-tuning language models on multi-hop document reasoning. Each example is a graded reasoning trace produced by a frontier AI agent solving a procedurally generated challenge from the Botcoin proof-of-inference network.
The traces contain no real domain knowledge. Entities are fictional, numbers are random, and documents are generated deterministically from 128-bit seeds. The reasoning structure is what matters:… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-causal-reasoning-tuning.causalverify-neurips2026
🎯 CausalVerify
An Execution-Grounded Benchmark for LLM Causal Inference Workflows
NeurIPS 2026 — Evaluations and Datasets Track · double-blind review · frozen at tag neurips2026-submission
💡 TL;DR
A benchmark of 259 published economics papers (Experiment A — real-paper text-agreement diagnostic) and 100 fixed-seed synthetic data-generating processes (Experiment B — execution-grounded coefficient recovery), evaluating 7 frontier LLMs. The central… See the full description on the dataset page: https://huggingface.co/datasets/causalverify/causalverify-neurips2026.task970_sherliic_causal_relationship
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task970_sherliic_causal_relationship
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task970_sherliic_causal_relationship.GPT-4-Self-Instruct-GermanHere we share a German dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied.
We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly German. This dataset will be updated continuously.
vi-gym-causal-ascii
Vi-Gym Causal ASCII Trajectories
This dataset contains autoregressive trajectories of a Large Language Model (LLM) agent learning spatial reasoning and geometric drawing within a simulated Vi (Vim) editor environment.
Warning
This dataset is a direct derivation of the source material, it might therefore also contain content not suitable for all audiences. All authors of the original artwork have full ownership.
Dataset Structure
Each record is a discrete step… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/vi-gym-causal-ascii.CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Anonymized release for double-blind review. Author, affiliation, and prior-whitepaper material have been removed. The data, solutions, and evaluation pipeline are otherwise identical to the version under review.
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections. The benchmark contains 173 queries over… See the full description on the dataset page: https://huggingface.co/datasets/anonsubmission16/CausalReasoningBenchmark.clt-tokenized-control-ar-zh-ko-ja-100m
Balanced Arabic-Chinese-Korean-Japanese 4B-token training data
Sequential FineWeb-2 documents tokenized with CausalNLP/gpt2-tokenizer-ar-zh-ko-ja. Each language contains at least 100,000,000 tokens in complete documents. Total target: 400,000,000 tokens. Approximately 100,000,000 tokens per Parquet shard.
Splits: arb_Arab, cmn_Hani, kor_Hang, jpn_Jpan. Schema: text: string, input_ids: list<int32>.
causal_kg_eval
Logic-Aware Causal Knowledge Graph Dataset
이 데이터셋은 비구조화된 텍스트에서 추출된 논리적 관계와 인과 경로의 타당성을 평가하기 위해 설계되었습니다. 단순한 유사도 기반 검색을 넘어, 구조화된 지식을 활용한 고난도 추론 성능 측정을 목적으로 합니다.
1. 개요 (Overview)
목적: 정보 간의 선후 관계, 인과성 및 다단계(Multi-hop) 연결성 검증
데이터 형식: JSON (head, relation, tail)
핵심 기능: Semantic Noise 필터링 및 논리적 추론 경로(Reasoning Path) 제공
2. 관계 스키마 (Relation Schema)
본 데이터셋은 정보 간의 연결 강도와 성격에 따라 다음 4가지 관계를 정의합니다.
Taxonomy (is_a): 상위 개념과 하위 개념 간의 계층적 분류
Causality (cause_of): 명확한 방향성을 가진… See the full description on the dataset page: https://huggingface.co/datasets/crjojo/causal_kg_eval.task392_inverse_causal_relationship
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task392_inverse_causal_relationship
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task392_inverse_causal_relationship.Refined-Anime-Text
Refined Anime Text for Continual Pre-training of Language Models
This is a subset of our novel synthetic dataset of anime-themed text, containing over 1M entries, ~440M GPT-4/3.5 tokens. This dataset has never been publicly released before. We are releasing this subset due to the community's interest in anime culture, which is underrepresented in general-purpose datasets, and the low quality of raw text due to the prevalence of internet slang and irrelevant content, making it… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Refined-Anime-Text.Retrieval-SFT-Chat
Retrieval-Based Multi-Turn Chat SFT Synthetic Data
A year ago, we released CausalLM/Refined-Anime-Text, a thematic subset of a dataset generated using the then state-of-the-art LLMs. This dataset comprises 1 million entries synthesized through long-context models that rewrote multi-document web text inputs, intended for continued pre-training. We are pleased to note that this data has been employed in various training scenarios and in studies concerning data and internet culture.
In… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Retrieval-SFT-Chat.causal_factors
Causal Factors Dataset
This dataset provides 'causal factors' for each sample in MMLU, BIG-Bench Hard (BBH) (specifically, the thirteen problem subset used in Turpin et al., 2023), and GPQA (specifically, GPQA Diamond). We use this dataset to test for the faithfulness and verbosity of models' chain of thought reasoning, to better understand how we can measure monitorability.
For each problem in these constituent datasets, we use a panel of judge models to extract any factors that… See the full description on the dataset page: https://huggingface.co/datasets/ameek/causal_factors.stride-preproc-climbmix
STRIDE: Preprocessed ClimbMix
Tokenized ClimbMix sequences used for STRIDE's pretraining-data attribution experiments. There is one training pool per released nanochat depth (d12, d16, d20, d24) and one held-out test set shared across depths. These are the corpora the released nanochat operators attribute over, so STRIDE's column indices line up with the lines of these files.
Files
File
Sequences
Size
Contents
climbmix_train_d12.jsonl
1,317,003
3.8 GB
training… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/stride-preproc-climbmix.clinical_causal_blindspot_probe_v0.1Clinical Causal Blindspot Probe
Detect when a clinician locks onto one cause and ignores alternative causal drivers.
Output JSON
blindspot
blindspot_type
correct_action
Run scoringpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
nam-causal-head-gating
NAM Causal Head Gating Datasets
Datasets for the nam-causal-head-gating Python package.
Paper: Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers (NeurIPS 2025)
Authors: Andrew Nam, Henry Conklin, Yukang Yang, Thomas Griffiths, Jonathan Cohen, Sarah-Jane Leslie
Datasets
aba_abb
Pattern recognition dataset for testing induction heads in transformer models.
Format: TSV (tab-separated values)
Columns: prompt, target… See the full description on the dataset page: https://huggingface.co/datasets/jonhanke-nam/nam-causal-head-gating.
