datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Compositional-ARCCompositional-ARC: Assessing Systematic Generalization in Abstract Spatial Reasoning
Philipp Mondorf, Shijia Zhou, Monica Riedler, and Barbara Plank. (2026). Compositional-ARC: Assessing systematic generalization in abstract spatial reasoning. In The Fourteenth International Conference on Learning Representations.
Systematic generalization refers to the capacity to understand and generate novel combinations from known components. Despite recent progress by large language… See the full description on the dataset page: https://huggingface.co/datasets/mainlp/Compositional-ARC.Composition-RL-EVA
Composition-RL
This repository contains the datasets presented in the paper Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models.
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that addresses the issue of "too-easy" prompts by automatically composing multiple verifiable problems into a single, more challenging yet still verifiable prompt. RL training on these compositional prompts helps… See the full description on the dataset page: https://huggingface.co/datasets/xx18/Composition-RL-EVA.CompositionalGSM_augmented
Compositional GSM_augmented
Compositional GSM_augmented is a math instruction dataset, inspired by Not All LLM Reasoners Are Created Equal.
It is based on nvidia/OpenMathInstruct-2 dataset, so you can use this dataset as training dataset.
It is generated using meta-llama/Meta-Llama-3.1-70B-Instruct model by Hyperbloic AI link. (Thanks for free credit!)
Replace the description of the data with the contents in the paper.
Each question in compositional GSM consists of two questions… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/CompositionalGSM_augmented.Physics-MATH-Composition-141K
Composition-RL
This repository contains the datasets for the paper Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models.
GitHub | Collection
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that combats the growing number of “too-easy” prompts (pass-rate = 1) by automatically composing multiple verifiable problems into a single, harder yet still-verifiable prompt. Across 4B–30B models… See the full description on the dataset page: https://huggingface.co/datasets/xx18/Physics-MATH-Composition-141K.MATH-Composition-199K
Composition-RL
Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
Code | Collection
Composition-RL is a data-efficient approach for Reinforcement Learning with Verifiable Rewards (RLVR). It addresses the issue of "too-easy" prompts (prompts that already achieve a pass rate of 1) by automatically composing multiple verifiable problems into a single, more challenging compositional prompt. This maintains informative training signals and… See the full description on the dataset page: https://huggingface.co/datasets/xx18/MATH-Composition-199K.MATH-Composition-Depth3
Composition-RL
Paper | Code | Collection
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that automatically composes multiple verifiable problems into a single, harder yet still-verifiable prompt. This method helps maintain informative training signals by combatting the growing number of "too-easy" prompts (pass-rate = 1) that occur during RL training.
Dataset Summary
This project introduces several compositional… See the full description on the dataset page: https://huggingface.co/datasets/xx18/MATH-Composition-Depth3.Polaris-Composition-1323K
Composition-RL Datasets
This repository contains datasets introduced in the paper Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models.
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that addresses the problem of "too-easy" prompts (pass-rate = 1) that occur during training. It automatically composes multiple verifiable problems into a single, harder verifiable prompt to maintain… See the full description on the dataset page: https://huggingface.co/datasets/xx18/Polaris-Composition-1323K.compositional-generalization-benchmark
Compositional Generalization Benchmark (CGB)
Benchmark accompanying "Beyond Benchmark Illusions: A Diagnostic Framework
for Compositional Generalization in LLM Mathematical Reasoning."
Overview
CGB tests whether LLM math reasoning generalizes across three types of
compositional perturbation applied to GSM8K problems: numerical
perturbation, structural reformulation, and clause injection. The
benchmark contains 1168 problems (300 source + 868 variants), evaluated… See the full description on the dataset page: https://huggingface.co/datasets/monanem/compositional-generalization-benchmark.RAG_Planning_Multi_Step_Composition
🇰🇿 Kazakh Multi-Step Planning and Tool Composition Dataset
Dataset Summary
Kazakh Multi-Step Planning and Tool Composition Dataset is a Kazakh-language dataset for training and evaluating Large Language Models (LLMs) in agentic AI workflows that require multi-step planning, tool composition, and structured function calling.
The dataset contains user requests, available tool schemas, expected tool calls, simulated tool responses, and complete multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/RAG_Planning_Multi_Step_Composition.
