datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NLR-Causal-Reasoning
SEA Causal Reasoning
SEA Causal Reasoning evaluates a model's ability to choose the correct cause or effect given a premise. It is sampled from XCOPA for Indonesian, Tamil, Thai, and Vietnamese.
Supported Tasks and Leaderboards
SEA Causal Reasoning is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Tamil (ta)
Thai (th)
Vietnamese (vi)… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLR-Causal-Reasoning.CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections (see CausalBenchmark.pdf). The benchmark contains 173 queries over 138 datasets. Each task is designed to evaluate both (i) identification, i.e., selecting an appropriate causal estimand and identification strategy given the study context, and (ii)… See the full description on the dataset page: https://huggingface.co/datasets/syrgkanislab/CausalReasoningBenchmark.Synthetic-Causal-Reasoning-50k
🏭 Sovereign Synthetic Reasoning Dataset (400k)
"High-Quality Chain-of-Thought Data at Scale."
📊 Overview
This dataset contains 400,000 synthetic reasoning samples spanning 16 enterprise domains (Finance, Pharma, Legal, Cybersecurity, Supply Chain, etc.).
It was generated using the Sovereign Generator, which produced 1.6 million samples and applied a strict quality filter (Top 25%) to retain only the most logically consistent and complex chains.
Average Quality… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/Synthetic-Causal-Reasoning-50k.domain-agnostic-causal-reasoning-tuning
Domain-Agnostic Causal Reasoning Tuning Dataset
Training data for fine-tuning language models on multi-hop document reasoning. Each example is a graded reasoning trace produced by a frontier AI agent solving a procedurally generated challenge from the Botcoin proof-of-inference network.
The traces contain no real domain knowledge. Entities are fictional, numbers are random, and documents are generated deterministically from 128-bit seeds. The reasoning structure is what matters:… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-causal-reasoning-tuning.causal-reasoning-enhenced
causal-reasoning-enhenced
Dataset Description
The Causal Reasoning Enhenced Dataset is designed to facilitate the understanding of causal relationships through structured reasoning tasks. This dataset features a collection of questions that challenge individuals to think critically about causality, providing detailed step-by-step reasoning for each scenario. It contains a variety of causal contexts, enhancing its applicability across different reasoning tasks. Key… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/causal-reasoning-enhenced.Causal-Reasoning-Assistant
Causal Reasoning Assistant (CRA)
A sophisticated AI assistant designed to provide expert-level causal inference analysis and reasoning. CRA leverages scientific methodologies and a structured approach to understand and explain causal relationships in various domains.
Overview
The Causal Reasoning Assistant (CRA) is built to offer rigorous and transparent causal analysis. It is designed to assist users in identifying causal relationships, formulating hypotheses… See the full description on the dataset page: https://huggingface.co/datasets/kojikubota/Causal-Reasoning-Assistant.causal-reasoning-benchmarks
Causal Reasoning Benchmarks
Datasets used in "On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning" (Deshmukh & Gupta, 2026).
Dataset Structure
train/transitivity_train.jsonl — 50,000 transitivity training examples
train/dsep_train.jsonl — 50,000 d-separation training examples
eval/length_eval.jsonl — 10,000 length generalization examples
eval/branching_eval.jsonl — 10,000 branching structure examples
eval/reversed_eval.jsonl — 10,000… See the full description on the dataset page: https://huggingface.co/datasets/ludwigw/causal-reasoning-benchmarks.CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Anonymized release for double-blind review. Author, affiliation, and prior-whitepaper material have been removed. The data, solutions, and evaluation pipeline are otherwise identical to the version under review.
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections. The benchmark contains 173 queries over… See the full description on the dataset page: https://huggingface.co/datasets/anonsubmission16/CausalReasoningBenchmark.causal_reasoningpara_Causal_Reasoningdagverse-examplestep_dpo_math_10k_causal_reasoningcausal-reasoning-ate
