datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omega-compositional
Compositional Math Problems
This dataset combines all compositional mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" with proper train/test splits. Each compositional setting includes training data from individual mathematical domains and test data consisting of compositional problems that require cross-domain reasoning.
Quick Start
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-compositional.compositionalitycomposition-classifications
NuBerea Composition Classifications
A curated reference set of scholarly-consensus composition history for the biblical corpus: the traditions behind the Old Testament, Deuterocanon, New Testament, and Old Testament Pseudepigrapha, and the source-critical relationships among them (e.g. Documentary Hypothesis strands, Markan priority, canonical collection, translation into the Septuagint). The dataset is a direct transcription of established scholarship — no machine learning or… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/composition-classifications.composition
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Charlie Zhang, Graham Neubig,
Xiang Yue
Carnegie Mellon University, Language Technologies Institute
Does Reinforcement Learning Truly Extend Reasoning?
This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/composition.compositionality_eccv_captioncompositionality_hpsv1compositionality_image_rewardcompositionality_seetruemulti-image-composition-instruction-following
Multi-Image Composition Instruction-Following
A large-scale multimodal dataset for multi-image composition via natural language instruction-following. Each case provides 2-3 input images (characters + scene) along with detailed Chinese instructions to compose them into a single photorealistic output image.
Designed for training and evaluating models on complex image composition tasks that require understanding of character identity preservation, pose generation, scene integration… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/multi-image-composition-instruction-following.Compositional-ARCCompositional-ARC: Assessing Systematic Generalization in Abstract Spatial Reasoning
Philipp Mondorf, Shijia Zhou, Monica Riedler, and Barbara Plank. (2026). Compositional-ARC: Assessing systematic generalization in abstract spatial reasoning. In The Fourteenth International Conference on Learning Representations.
Systematic generalization refers to the capacity to understand and generate novel combinations from known components. Despite recent progress by large language… See the full description on the dataset page: https://huggingface.co/datasets/mainlp/Compositional-ARC.Food-Composition
Ingredients CSV/Parquet File
Overview
The following data comes from the United States Department of Agriculture’s Food Composition Database. It contains data for various types of food ingredients including the amounts of different vitamins and minerals found in the foods as well as macronutrient percentages. The food covered spans a large variety of foods from butter to Campbell’s soup. Much of the supplementary documenation for each field comes directly from that pages’… See the full description on the dataset page: https://huggingface.co/datasets/hootan09/Food-Composition.mars-chemcam-compositions
Mars ChemCam LIBS Oxide Compositions
Part of the Planetary Science Datasets collection on Hugging Face.
Major oxide compositions of Mars surface rock and soil targets analyzed by the
Chemistry and Camera (ChemCam) Laser-Induced Breakdown Spectroscopy (LIBS)
instrument aboard the Curiosity rover. Currently 30,458 individual
point analyses across 4,184 named targets, spanning sols
0 to 4612.
Dataset description
ChemCam fires a focused laser pulse at rock and soil… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/mars-chemcam-compositions.Composition-RL-EVA
Composition-RL
This repository contains the datasets presented in the paper Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models.
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that addresses the issue of "too-easy" prompts by automatically composing multiple verifiable problems into a single, more challenging yet still verifiable prompt. RL training on these compositional prompts helps… See the full description on the dataset page: https://huggingface.co/datasets/xx18/Composition-RL-EVA.composition-10B-rlCompositionalGSM_augmented
Compositional GSM_augmented
Compositional GSM_augmented is a math instruction dataset, inspired by Not All LLM Reasoners Are Created Equal.
It is based on nvidia/OpenMathInstruct-2 dataset, so you can use this dataset as training dataset.
It is generated using meta-llama/Meta-Llama-3.1-70B-Instruct model by Hyperbloic AI link. (Thanks for free credit!)
Replace the description of the data with the contents in the paper.
Each question in compositional GSM consists of two questions… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/CompositionalGSM_augmented.compositionality-subsample
Dataset Card for "compositionality-subsample"
More Information needed
Physics-MATH-Composition-141K
Composition-RL
This repository contains the datasets for the paper Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models.
GitHub | Collection
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that combats the growing number of “too-easy” prompts (pass-rate = 1) by automatically composing multiple verifiable problems into a single, harder yet still-verifiable prompt. Across 4B–30B models… See the full description on the dataset page: https://huggingface.co/datasets/xx18/Physics-MATH-Composition-141K.vlm-compositionality-embeddings
VLM Compositionality Embeddings
Pre-computed image and text embeddings for the thesis "From Euclidean to Hyperbolic Vision-Language Spaces: A Study of Attribute–Object Compositionality" by Meelad Dashti (Politecnico di Torino & University of Twente, 2026).
Code repository: github.com/MelDashti/hyperbolic-vlm-compositionality
Models
Model
Geometry
Architecture
Training Data
CLIP ViT-L/14
Spherical
ViT-L/14
WIT (400M+ pairs)
DINOv2 ViT-L/14
Spherical
ViT-L/14… See the full description on the dataset page: https://huggingface.co/datasets/Meldashti/vlm-compositionality-embeddings.ssa-body-composition-women
SSA Body Composition Dataset (Women, Multi-ancestry) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: other_unclassified - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/ssa-body-composition-women.food-composition-matrix
Food Composition Nutrient Matrix — TKPI 2017 & USDA SR Legacy 2018
This repository contains two food composition datasets reformatted as wide-format nutrient matrices, suitable for a wide range of research tasks including nutrient prediction, food type classification, missing value imputation, dietary analysis, and other machine learning applications on food data. Both datasets share a harmonised set of 18 common nutrients, enabling cross-dataset generalization experiments.… See the full description on the dataset page: https://huggingface.co/datasets/ULM-DS-Lab/food-composition-matrix.digital-twin-composition
Digital Twin Composition
Datasets for retrieving and filling DTDL
digital-twin interfaces from natural-language requests. All parts live in this one repo
as separate configs.
Two families
The configs come in two provenances that share the same schemas but must not be mixed:
synthetic (triplets, interfaces, fill_eval, topics, eval_small, eval_mid)
— LLM-generated interfaces and everything derived from them. This is the training pool.
real (interfaces_real… See the full description on the dataset page: https://huggingface.co/datasets/zirenx/digital-twin-composition.composition-10B-testMATH-Composition-199K
Composition-RL
Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
Code | Collection
Composition-RL is a data-efficient approach for Reinforcement Learning with Verifiable Rewards (RLVR). It addresses the issue of "too-easy" prompts (prompts that already achieve a pass rate of 1) by automatically composing multiple verifiable problems into a single, more challenging compositional prompt. This maintains informative training signals and… See the full description on the dataset page: https://huggingface.co/datasets/xx18/MATH-Composition-199K.MATH-Composition-Depth3
Composition-RL
Paper | Code | Collection
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that automatically composes multiple verifiable problems into a single, harder yet still-verifiable prompt. This method helps maintain informative training signals by combatting the growing number of "too-easy" prompts (pass-rate = 1) that occur during RL training.
Dataset Summary
This project introduces several compositional… See the full description on the dataset page: https://huggingface.co/datasets/xx18/MATH-Composition-Depth3.compositional-safety-folds
Compositional Safety Policy Benchmark — Contrastive Folds
Dataset Summary
This dataset evaluates whether language models apply written safety policies
compositionally, as opposed to responding to lexical features of a request. Each
instance pairs a self-contained policy of seven or eight numbered rules with a
user request, and is labelled with the action the policy requires and the subset
of rules that determine it.
Instances are organised into contrastive folds:… See the full description on the dataset page: https://huggingface.co/datasets/zmsy/compositional-safety-folds.glass_alloy_compositionThis is an alloy composition datasetasia-owid-dietary-composition-by-country
Dietary Composition By Country | Asia (Our World in Data)
🌏 2,441 observations · 47 Asia countries · 1961–2023 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 2,441 observations of Dietary Composition By Country data across 47 Asia countries, spanning 1961–2023.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Dietary Composition By Country
Geographic coverage
47… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-dietary-composition-by-country.compositionality_aigciqa2023Polaris-Composition-1323K
Composition-RL Datasets
This repository contains datasets introduced in the paper Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models.
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that addresses the problem of "too-easy" prompts (pass-rate = 1) that occur during training. It automatically composes multiple verifiable problems into a single, harder verifiable prompt to maintain… See the full description on the dataset page: https://huggingface.co/datasets/xx18/Polaris-Composition-1323K.dyt-composition-artifacts
DyT Composition Study Artifacts
This dataset contains sanitized result manifests and analysis outputs for
When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer.
DOI: https://doi.org/10.48550/arXiv.2604.23434
Contents
The artifacts include aggregate training metrics, saturation measurements, statistical-test summaries, predictor-validation outputs, table-source manifests, and selected aggregate analysis files used by the… See the full description on the dataset page: https://huggingface.co/datasets/lucky-verma/dyt-composition-artifacts.
