datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
composition
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Charlie Zhang, Graham Neubig,
Xiang Yue
Carnegie Mellon University, Language Technologies Institute
Does Reinforcement Learning Truly Extend Reasoning?
This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/composition.compositional_causal_reasoning
– 3k+ Hugging Face downloads –
https://jmaasch.github.io/ccr/
Causal reasoning and compositional reasoning are two core aspirations in AI. Measuring these behaviors requires principled
evaluation methods. Maasch et al. (2025) consider both behaviors simultaneously, under
the umbrella of compositional causal reasoning (CCR): the ability to infer how causal measures compose and, equivalently, how causal quantities propagate
through graphs. CCR.GB applies the… See the full description on the dataset page: https://huggingface.co/datasets/jmaasch/compositional_causal_reasoning.CompositionalGSM_augmented
Compositional GSM_augmented
Compositional GSM_augmented is a math instruction dataset, inspired by Not All LLM Reasoners Are Created Equal.
It is based on nvidia/OpenMathInstruct-2 dataset, so you can use this dataset as training dataset.
It is generated using meta-llama/Meta-Llama-3.1-70B-Instruct model by Hyperbloic AI link. (Thanks for free credit!)
Replace the description of the data with the contents in the paper.
Each question in compositional GSM consists of two questions… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/CompositionalGSM_augmented.nanochat-depo-composition-depth2-w4-retry-20260715
Nanochat Depo composition v1
Each 16-node single-cycle graph yields eight independent one-query documents:
four base starts paired across query depths (1, 2).
This source contains train and validation splits only. Phase depth is
2; the materialized context width is 4.
nanochat-depo-composition-depth2-w4-pubfix-20260715
Nanochat Depo composition v1
Each 16-node single-cycle graph yields eight independent one-query documents:
four base starts paired across query depths (1, 2).
This source contains train and validation splits only. Phase depth is
2; the materialized context width is 4.
nanochat-depo-composition-depth2-w4-20260715
Nanochat Depo composition v1
Each 16-node single-cycle graph yields eight independent one-query documents:
four base starts paired across query depths (1, 2).
This source contains train and validation splits only. Phase depth is
2; the materialized context width is 4.
compositional-generalization-benchmark
Compositional Generalization Benchmark (CGB)
Benchmark accompanying "Beyond Benchmark Illusions: A Diagnostic Framework
for Compositional Generalization in LLM Mathematical Reasoning."
Overview
CGB tests whether LLM math reasoning generalizes across three types of
compositional perturbation applied to GSM8K problems: numerical
perturbation, structural reformulation, and clause injection. The
benchmark contains 1168 problems (300 source + 868 variants), evaluated… See the full description on the dataset page: https://huggingface.co/datasets/monanem/compositional-generalization-benchmark.nanochat-depo-composition-depth2-w4-transport-20260715
Nanochat Depo composition v1
Each 16-node single-cycle graph yields eight independent one-query documents:
four base starts paired across query depths (1, 2).
This source contains train and validation splits only. Phase depth is
2; the materialized context width is 4.
RAG_Planning_Multi_Step_Composition
🇰🇿 Kazakh Multi-Step Planning and Tool Composition Dataset
Dataset Summary
Kazakh Multi-Step Planning and Tool Composition Dataset is a Kazakh-language dataset for training and evaluating Large Language Models (LLMs) in agentic AI workflows that require multi-step planning, tool composition, and structured function calling.
The dataset contains user requests, available tool schemas, expected tool calls, simulated tool responses, and complete multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/RAG_Planning_Multi_Step_Composition.
