causal-reasoning
NLR-Causal-Reasoning
SEA Causal Reasoning
SEA Causal Reasoning evaluates a model's ability to choose the correct cause or effect given a premise. It is sampled from XCOPA for Indonesian, Tamil, Thai, and Vietnamese.
Supported Tasks and Leaderboards
SEA Causal Reasoning is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Tamil (ta)
Thai (th)
Vietnamese (vi)… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLR-Causal-Reasoning.CausalReasoningBenchmark
Automated Causal Reasoning Benchmark
Overview
The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections (see CausalBenchmark.pdf). The benchmark contains 173 queries over 138 datasets. Each task is designed to evaluate both (i) identification, i.e., selecting an appropriate causal estimand and identification strategy given the study context, and (ii)… See the full description on the dataset page: https://huggingface.co/datasets/syrgkanislab/CausalReasoningBenchmark.Synthetic-Causal-Reasoning-50k
🏭 Sovereign Synthetic Reasoning Dataset (400k)
"High-Quality Chain-of-Thought Data at Scale."
📊 Overview
This dataset contains 400,000 synthetic reasoning samples spanning 16 enterprise domains (Finance, Pharma, Legal, Cybersecurity, Supply Chain, etc.).
It was generated using the Sovereign Generator, which produced 1.6 million samples and applied a strict quality filter (Top 25%) to retain only the most logically consistent and complex chains.
Average Quality… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/Synthetic-Causal-Reasoning-50k.Causal-Reasoning-Bench_CRBench
🦙 Causal Reasoning Benchmark (CRBench)
CRBench is a benchmark for evaluating process-level causal failures in
Chain-of-Thought (CoT) reasoning.
Rather than treating incorrect reasoning traces as homogeneous failures,
CRBench characterizes erroneous dependencies among intermediate reasoning
steps through a step-level causal-error taxonomy. It is designed to evaluate
whether reasoning methods can identify and correct structured causal failures
that arise during the reasoning… See the full description on the dataset page: https://huggingface.co/datasets/EdmondFU/Causal-Reasoning-Bench_CRBench.Ordis-CausalReasoning-92K-Verified
Ordis CausalReasoning 92K Verified
This is NOT another text dataset.
Each entry is a complete civilization simulation with 20,000+ structured data points, verified causal graphs, and ground truth outcomes.
What is Ordis?
Ordis is a proprietary multi-agent simulation engine (the "Liquid Universe Engine") that runs thousands of autonomous agents through 5,000-tick lifecycles. Agents gather resources, share, steal, betray, form coalitions, discover emergent strategies (DVS… See the full description on the dataset page: https://huggingface.co/datasets/sugiken/Ordis-CausalReasoning-92K-Verified.domain-agnostic-causal-reasoning-tuning
Domain-Agnostic Causal Reasoning Tuning Dataset
Training data for fine-tuning language models on multi-hop document reasoning. Each example is a graded reasoning trace produced by a frontier AI agent solving a procedurally generated challenge from the Botcoin proof-of-inference network.
The traces contain no real domain knowledge. Entities are fictional, numbers are random, and documents are generated deterministically from 128-bit seeds. The reasoning structure is what matters:… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-causal-reasoning-tuning.
