datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
composition
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Charlie Zhang, Graham Neubig,
Xiang Yue
Carnegie Mellon University, Language Technologies Institute
Does Reinforcement Learning Truly Extend Reasoning?
This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/composition.multi-image-composition-instruction-following
Multi-Image Composition Instruction-Following
A large-scale multimodal dataset for multi-image composition via natural language instruction-following. Each case provides 2-3 input images (characters + scene) along with detailed Chinese instructions to compose them into a single photorealistic output image.
Designed for training and evaluating models on complex image composition tasks that require understanding of character identity preservation, pose generation, scene integration… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/multi-image-composition-instruction-following.Compositional-ARCCompositional-ARC: Assessing Systematic Generalization in Abstract Spatial Reasoning
Philipp Mondorf, Shijia Zhou, Monica Riedler, and Barbara Plank. (2026). Compositional-ARC: Assessing systematic generalization in abstract spatial reasoning. In The Fourteenth International Conference on Learning Representations.
Systematic generalization refers to the capacity to understand and generate novel combinations from known components. Despite recent progress by large language… See the full description on the dataset page: https://huggingface.co/datasets/mainlp/Compositional-ARC.composition-10B-rlcomposition-10B-testdyt-composition-artifacts
DyT Composition Study Artifacts
This dataset contains sanitized result manifests and analysis outputs for
When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer.
DOI: https://doi.org/10.48550/arXiv.2604.23434
Contents
The artifacts include aggregate training metrics, saturation measurements, statistical-test summaries, predictor-validation outputs, table-source manifests, and selected aggregate analysis files used by the… See the full description on the dataset page: https://huggingface.co/datasets/lucky-verma/dyt-composition-artifacts.composition-10B-valskill_composition_hypothesis
A Dataset for Skill Composition Hypothesis
We employ a language model to annotate the specific skills assessed by each question derived from various Natural Language Processing benchmarks. The skill taxonomy utilized is sourced from IXL. The associated GitHub repository that produces this dataset can be found here: https://github.com/sangttruong/skill-composition-hypothesis.
compositional_gsmsand_composition.jsontimex-compositional-sentence
