datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DAPO-Math-17kDAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill
DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill
A high-quality Chain-of-Thought (CoT) dataset generated using Qwen/Qwen3-235B-A22B-Thinking-2507 with rejection sampling on BytedTsinghua-SIA/DAPO-Math-17k. This dataset is ideal for SFT distillation training to improve mathematical reasoning capabilities of models.
The dataset format is compatible with LLaMA-Factory for efficient SFT training.
Files
dapo_distill_boxed.json: Single sampling subset (15,129… See the full description on the dataset page: https://huggingface.co/datasets/Yang-Zhou/DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill.dapo-math-17k-verl
DAPO-Math-17K VERL
📊 Dataset Summary
This dataset contains 17,147 mathematical reasoning problems in VERL format, processed from haizhongzheng/DAPO-Math-17K-cleaned.
Key Features:
17,147 high-quality math problems
Converted to VERL format for reward modeling
Verified ground truth answers
Ready for reinforcement learning training
🔗 Source Dataset
Original Repository
Repository: haizhongzheng/DAPO-Math-17K-cleaned
License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/dapo-math-17k-verl.Staleness-GRPO-DAPO-Math-17k
Staleness GRPO DAPO Math 17k
The exact 17,005-row training dataset shared by the staleness-cap-2 Qwen2.5-Math-1.5B, Qwen2.5-3B, and Qwen2.5-Math-7B checkpoints, and the staleness-cap-4 Qwen2.5-Math-1.5B checkpoint. All four training manifests record the same SHA-256 for the training file.
Source and processing
Derived from the all configuration of open-r1/DAPO-Math-17k-Processed, itself processed from BytedTsinghua-SIA/DAPO-Math-17k. Source revision:… See the full description on the dataset page: https://huggingface.co/datasets/zbeeb/Staleness-GRPO-DAPO-Math-17k.DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-non-thinking-dedup
DAPO Math Qwen3-235B non-thinking, deduplicated
This dataset is derived from
Yang-Zhou/DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill, specifically
dapo_distill_boxed_non_thinking.json.
It retains the original LLaMA-Factory-compatible instruction, input, and
output columns. Duplicate rows are identified by collapsing consecutive
whitespace in instruction, trimming leading/trailing whitespace, and hashing
the normalized instruction. The first row in each duplicate… See the full description on the dataset page: https://huggingface.co/datasets/YYYYYYibo/DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-non-thinking-dedup.dapo-math-17k-qwen3-1.7b-base-n8
DAPO-Math-17k sampled with Qwen3-1.7B-Base, n=8
17398 problems from the RL training set, each sampled 8 times and scored with
the reward function the RL runs themselves used.
The point of this dataset is to be comparable with what the RL runs actually saw, so
every sampling knob is taken from the live training config or from the default that
config falls through to. Two of them are not in the config file at all and would be
wrong if guessed: top_k = -1 and min_tokens = 1.… See the full description on the dataset page: https://huggingface.co/datasets/RyanYr/dapo-math-17k-qwen3-1.7b-base-n8.DAPO-Math-17kdapo-math-giant-tree
DAPO-Math Giant Trees (Qwen2.5-Math-7B)
Per-prompt deep search trees generated with Qwen2.5-Math-7B over prompts from
the DAPO-Math-15k training set. Each tree is a chunked, branching rollout used for
analysis / data distillation (not RL training).
Generation
Model: Qwen2.5-Math-7B, temperature 1.0, top_p 1.0
Chunk size: 128 tokens per generation step
Branching factors by depth: [4, 16, 16, 16] (depths 0–3 branch; deeper depths continue with bf=1)
Max depth: 32… See the full description on the dataset page: https://huggingface.co/datasets/Jianshu001/dapo-math-giant-tree.dapo-filtered
dapo-filtered
A small, difficulty-banded slice of competition math, cut so that a specific policy
solves each problem rarely but not never. Built for rejection fine-tuning (RFT) and
RL experiments, where a set the model never solves gives nothing to bootstrap from and a
set it always solves gives nothing to learn.
Source: open-r1/DAPO-Math-17k-Processed
(config en). Answers are bare integers; the reward is exact-match on a \boxed{} answer.
Files
file
rows… See the full description on the dataset page: https://huggingface.co/datasets/kushasareen/dapo-filtered.DAPO-Math-17k-MATH-500DAPO-Math-17k with the MATH-500 test split converted to the same parquet schema and prompt format.
DAPO-Math-Multilingual-6Lang
DAPO-Math 17k · Multilingual (6 languages) — RL prompts
Verifiable-reward math RL prompts in six UN languages (English, Chinese, Spanish, French, Arabic,
Russian), balanced and round-robin interleaved for GRPO. Each problem is posed in one target
language with an instruction to reason entirely in that language, plus a rule-based ground-truth
answer — for training and studying language-consistent multilingual reasoning (models that reason
in the target language rather than… See the full description on the dataset page: https://huggingface.co/datasets/96kevinli29/DAPO-Math-Multilingual-6Lang.DAPO-17K-Plus
RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data
Introduction
DAPO-17k-Plus (DAPO++) is the dataset presented in the paper RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data.
Acknowledgements
DAPO++ is built on the following repositories and we thank their teams for their valuable contributions to the community:
DAPO
Citation
If you find our work useful, feel… See the full description on the dataset page: https://huggingface.co/datasets/CelineHuangxy/DAPO-17K-Plus.math_merged_deduped_OR1_dapo
Math subset for training L1 using RL
This dataset is inspired by LLM360/Reasoning360(GURU92K-math), but reproduced from DAPO-Math-17K and Skywork-OR1-Math. DeepScaleR was not used for source duplications.
Dataset Details
Dataset Description
Curated by: Leon (Me)
Funded by [optional]: AIGCode/Koting Intelligence
Language(s) (NLP): Mostly in English with a few in Chinese
License: MIT (following GURU-92K)
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/Leon-Leee/math_merged_deduped_OR1_dapo.DAPO-Math-Unique-17k
[!NOTE]
This version is deduplicated according to the "raw_problem_id" column, i.e., the hash value of the raw problem strings.
The user message content follows the template below:
"""\
Solve the following math problem step by step. The last line of your response should be of the form Answer: $Answer (without quotes) where $Answer is the answer to the problem.
{raw_problem}
Remember to put your answer on its own line after "Answer:".\
"""
DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4
DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 128,000
Unique prompts: 32,000
Responses per prompt: 4
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and teacher… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4.DAPO-Math-Raw-17k
[!NOTE]
This version is deduplicated according to the "index" in the "extra_info" column, i.e., the original UUID assigned in the raw dataset before duplication.
DAPO-Gemma3-27B-IT-RL-SFT-Data-correct
DAPO-Gemma3-27B-IT-RL-SFT-Data-correct
Filtered subset of
JWei05/DAPO-Gemma3-27B-IT-RL-SFT-Data:
only the teacher responses whose final answer is math_verify-correct against
the original DAPO-Math-17k ground truth.
Stats
Source rows: 69,592 (17,398 prompts × 4 teacher responses)
Kept rows: 41,831 (60.1%)
Prompts with ≥1 correct response: 13,062 / 17,398 (75.1%)
Prompts with 4/4 correct responses: 7,492 (43.1%)
Scoring
Same function as used during RL… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-IT-RL-SFT-Data-correct.DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4
DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 133,184
Unique prompts: 33,296
Responses per prompt: 4
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4.DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data
DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 66,592
Unique prompts: 33,296
Responses per prompt: 2
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and teacher assistant… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data.DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data-all33296-n4
DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data-all33296-n4
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-12b-pt-from-step60-seed43, subfolder step_000020
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 133,184
Unique prompts: 33,296
Responses per prompt: 4
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data-all33296-n4.autoteacher-dapo-claude-solved-coded
autoteacher-dapo-claude-solved-coded
zjhhhh/autoteacher-dapo-claude-solved (512 DAPO math problems with Claude-written, student-verified hints) augmented with a strategy coding of every hint against a compact codebook of 74 reusable problem-solving strategies.
Companion codebook dataset: zjhhhh/autoteacher-dapo-codebook.
Added columns
column
type
description
hint_code_indices
list[int]
code_ids (1–74) of the codebook strategies that form the core of the… See the full description on the dataset page: https://huggingface.co/datasets/zjhhhh/autoteacher-dapo-claude-solved-coded.DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data
DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-12b-pt-from-step60-seed43, subfolder step_000020
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 66,592
Unique prompts: 33,296
Responses per prompt: 2
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and teacher assistant… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data.ptdbench-data-format-task-agent-loop-022-llama-dapo-math-dataset
PTDBench dataset snapshot: task_agent_loop_022-llama-dapo-math
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: data_format
Source evaluation metric: val-core/math_dapo/reward/mean@1
Provenance: Processed from BytedTsinghua-SIA/DAPO-Math-17k; task-specific bytes are pinned.
License: Apache-2.0
The artifact manifest records every hydrated runtime… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-data-format-task-agent-loop-022-llama-dapo-math-dataset.DAPO-Gemma3-1B-PT-DAPO-17.4k
DAPO-Gemma3-1B-PT-DAPO-17.4k
Traces sampled from google/gemma-3-1b-pt on the DAPO-Math-17k train + 100-question val splits,
using the SAME unified few-shot chat prompt and sampling (temp 1.0, top_p 1.0, top_k -1, 20k max,
single BOS) as RL training. 16 samples per question. Splits: train (17,198 q), validation (100 q).
Columns: prompt_text, response_text, prompt_token_ids, response_token_ids, input_ids, response_mask,
teacher_log_probs, prompt_idx (shared across a question's 16… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-1B-PT-DAPO-17.4k.DAPO-Math-17kDAPO-Math-8k-Stratified
DAPO-Math-8k-Stratified
A fixed 8,000-problem stratified random subset of BytedTsinghua-SIA/DAPO-Math-17k for efficient RLVR experimentation.
Creation
Source: DAPO-Math-17k (17,917 unique problems after deduplication)
Stratification: 5 strata by prompt length (quintiles), proportional sampling
Random seed: 42
Split: 7,500 train / 500 validation
Distribution Match
The subset preserves the prompt length distribution of the full dataset:
Percentile
Full… See the full description on the dataset page: https://huggingface.co/datasets/eshwarprasadS/DAPO-Math-8k-Stratified.dapo-en-10k
Understanding Tool-Integrated Reasoning Training Dataset
This is the training dataset for the paper Understanding Tool-Integrated Reasoning.
This dataset is randomly sampled from DAPO dataset, used to study why Tool-Integrated Reasoning (TIR) makes Large Language Models (LLMs) more capable.
ptdbench-llama-dapo-implementation-task-monkey-patch-011-dataset
PTDBench dataset snapshot: task_monkey_patch_011
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: llama_dapo_implementation
Source evaluation metric: val-core/math_dapo/acc/mean@1
Provenance: Processed from BytedTsinghua-SIA/DAPO-Math-17k; task-specific bytes are pinned.
License: Apache-2.0
The artifact manifest records every hydrated runtime path… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-llama-dapo-implementation-task-monkey-patch-011-dataset.DAPO-OpenMathInstruct2-34k
DAPO + OpenMathInstruct-2 Mix (34k)
A 50/50 mix of two math-reasoning datasets used for RL training of Gemma 3 PT models with DAPO (GRPO).
Composition
Source
Rows
Description
open-r1/DAPO-Math-17k-Processed
17,398
DAPO training set (AoPS + competition math)
nvidia/OpenMathInstruct-2 subset
17,398
Synthetic augmented math problems
Total
34,796
Within the OpenMathInstruct-2 subset:
14,529 augmented_math (competition-style augmentations)
2,372… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-OpenMathInstruct2-34k.DAPO-Math-17k
