datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omega-compositional
Compositional Math Problems
This dataset combines all compositional mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" with proper train/test splits. Each compositional setting includes training data from individual mathematical domains and test data consisting of compositional problems that require cross-domain reasoning.
Quick Start
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-compositional.compositionalitycomposition-classifications
NuBerea Composition Classifications
A curated reference set of scholarly-consensus composition history for the biblical corpus: the traditions behind the Old Testament, Deuterocanon, New Testament, and Old Testament Pseudepigrapha, and the source-critical relationships among them (e.g. Documentary Hypothesis strands, Markan priority, canonical collection, translation into the Septuagint). The dataset is a direct transcription of established scholarship — no machine learning or… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/composition-classifications.compositionality_hpsv1compositionality_eccv_captioncompositionality_image_rewardComposition-RL-EVA
Composition-RL
This repository contains the datasets presented in the paper Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models.
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that addresses the issue of "too-easy" prompts by automatically composing multiple verifiable problems into a single, more challenging yet still verifiable prompt. RL training on these compositional prompts helps… See the full description on the dataset page: https://huggingface.co/datasets/xx18/Composition-RL-EVA.compositionality_seetruemars-chemcam-compositions
Mars ChemCam LIBS Oxide Compositions
Part of the Planetary Science Datasets collection on Hugging Face.
Major oxide compositions of Mars surface rock and soil targets analyzed by the
Chemistry and Camera (ChemCam) Laser-Induced Breakdown Spectroscopy (LIBS)
instrument aboard the Curiosity rover. Currently 30,458 individual
point analyses across 4,184 named targets, spanning sols
0 to 4612.
Dataset description
ChemCam fires a focused laser pulse at rock and soil… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/mars-chemcam-compositions.CompositionalGSM_augmented
Compositional GSM_augmented
Compositional GSM_augmented is a math instruction dataset, inspired by Not All LLM Reasoners Are Created Equal.
It is based on nvidia/OpenMathInstruct-2 dataset, so you can use this dataset as training dataset.
It is generated using meta-llama/Meta-Llama-3.1-70B-Instruct model by Hyperbloic AI link. (Thanks for free credit!)
Replace the description of the data with the contents in the paper.
Each question in compositional GSM consists of two questions… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/CompositionalGSM_augmented.Physics-MATH-Composition-141K
Composition-RL
This repository contains the datasets for the paper Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models.
GitHub | Collection
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that combats the growing number of “too-easy” prompts (pass-rate = 1) by automatically composing multiple verifiable problems into a single, harder yet still-verifiable prompt. Across 4B–30B models… See the full description on the dataset page: https://huggingface.co/datasets/xx18/Physics-MATH-Composition-141K.compositionality-subsample
Dataset Card for "compositionality-subsample"
More Information needed
RL-Compositionality-Stage1-RFT-DataStage 1 RFT data.
Paper: https://huggingface.co/papers/2509.25123
Code: https://github.com/PRIME-RL/RL-Compositionality
MATH-Composition-199K
Composition-RL
Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
Code | Collection
Composition-RL is a data-efficient approach for Reinforcement Learning with Verifiable Rewards (RLVR). It addresses the issue of "too-easy" prompts (prompts that already achieve a pass rate of 1) by automatically composing multiple verifiable problems into a single, more challenging compositional prompt. This maintains informative training signals and… See the full description on the dataset page: https://huggingface.co/datasets/xx18/MATH-Composition-199K.Polaris-Composition-1323K
Composition-RL Datasets
This repository contains datasets introduced in the paper Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models.
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that addresses the problem of "too-easy" prompts (pass-rate = 1) that occur during training. It automatically composes multiple verifiable problems into a single, harder verifiable prompt to maintain… See the full description on the dataset page: https://huggingface.co/datasets/xx18/Polaris-Composition-1323K.compositional-safety-folds
Compositional Safety Policy Benchmark — Contrastive Folds
Dataset Summary
This dataset evaluates whether language models apply written safety policies
compositionally, as opposed to responding to lexical features of a request. Each
instance pairs a self-contained policy of seven or eight numbered rules with a
user request, and is labelled with the action the policy requires and the subset
of rules that determine it.
Instances are organised into contrastive folds:… See the full description on the dataset page: https://huggingface.co/datasets/zmsy/compositional-safety-folds.digital-twin-composition
Digital Twin Composition
Datasets for retrieving and filling DTDL
digital-twin interfaces from natural-language requests. All parts live in this one repo
as separate configs.
Two families
The configs come in two provenances that share the same schemas but must not be mixed:
synthetic (triplets, interfaces, fill_eval, topics, eval_small, eval_mid)
— LLM-generated interfaces and everything derived from them. This is the training pool.
real (interfaces_real… See the full description on the dataset page: https://huggingface.co/datasets/zirenx/digital-twin-composition.permutation-compositionssa-body-composition-women
SSA Body Composition Dataset (Women, Multi-ancestry) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: other_unclassified - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/ssa-body-composition-women.MATH-Composition-Depth3
Composition-RL
Paper | Code | Collection
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that automatically composes multiple verifiable problems into a single, harder yet still-verifiable prompt. This method helps maintain informative training signals by combatting the growing number of "too-easy" prompts (pass-rate = 1) that occur during RL training.
Dataset Summary
This project introduces several compositional… See the full description on the dataset page: https://huggingface.co/datasets/xx18/MATH-Composition-Depth3.RL-Compositionality-Stage2-RL-Level8-TestDataStage 2 RL Level 1 to 8 evaluation data.
Paper: https://huggingface.co/papers/2509.25123
Code: https://github.com/PRIME-RL/RL-Compositionality
asia-owid-dietary-composition-by-country
Dietary Composition By Country | Asia (Our World in Data)
🌏 2,441 observations · 47 Asia countries · 1961–2023 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 2,441 observations of Dietary Composition By Country data across 47 Asia countries, spanning 1961–2023.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Dietary Composition By Country
Geographic coverage
47… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-dietary-composition-by-country.compositionality_aigciqa2023RL-Compositionality-Stage2-RL-Level1-TrainDataStage 2 RL Level 1 data.
Paper: https://huggingface.co/papers/2509.25123
Code: https://github.com/PRIME-RL/RL-Compositionality
terminal_bench_2_a1_multifile_composition_20260326_040008RL-Compositionality-Stage2-RL-Level2-TrainDataStage 2 RL Level 2 data.
Paper: https://huggingface.co/papers/2509.25123
Code: https://github.com/PRIME-RL/RL-Compositionality
drugs-composition-indonesian-donut
Dataset Card for "drugs-composition-indonesian-donut"
Generate Custom Data
Please visit https://huggingface.co/spaces/jonathanjordan21/donut-labelling for the interface to generate custom data.
The data format is (.zip). Images and Labels are stored in separated .zip files.
[More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards
asia-owid-per-capita-diet-composition
Per Capita Diet Composition | Asia (Our World in Data)
🌏 2,441 observations · 47 Asia countries · 1961–2023 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 2,441 observations of Per Capita Diet Composition data across 47 Asia countries, spanning 1961–2023.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Per Capita Diet Composition
Geographic coverage
47 Asia… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-per-capita-diet-composition.terminal_bench_2_a1_multifile_composition_20260324_100610africa-tunisia-composition-foret-oleicole-sfax-6e43c362
Composition Foret Oleicole Sfax | Africa (Tunisia Open Data)
14 rows - 1 Africa country/area - 2017-2018 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 14 rows from Tunisia Open Data, covering Composition Foret Oleicole Sfax. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures
Agriculture datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-tunisia-composition-foret-oleicole-sfax-6e43c362.
