datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Synthetic-Causal-Reasoning-50k
🏭 Sovereign Synthetic Reasoning Dataset (400k)
"High-Quality Chain-of-Thought Data at Scale."
📊 Overview
This dataset contains 400,000 synthetic reasoning samples spanning 16 enterprise domains (Finance, Pharma, Legal, Cybersecurity, Supply Chain, etc.).
It was generated using the Sovereign Generator, which produced 1.6 million samples and applied a strict quality filter (Top 25%) to retain only the most logically consistent and complex chains.
Average Quality… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/Synthetic-Causal-Reasoning-50k.causal-reasoning-benchmarks
Causal Reasoning Benchmarks
Datasets used in "On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning" (Deshmukh & Gupta, 2026).
Dataset Structure
train/transitivity_train.jsonl — 50,000 transitivity training examples
train/dsep_train.jsonl — 50,000 d-separation training examples
eval/length_eval.jsonl — 10,000 length generalization examples
eval/branching_eval.jsonl — 10,000 branching structure examples
eval/reversed_eval.jsonl — 10,000… See the full description on the dataset page: https://huggingface.co/datasets/ludwigw/causal-reasoning-benchmarks.
