datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DAPO-Math-17kdapo-math-17kDAPO-Math-17k-Processed
Dataset Card for DAPO-Math-17k-Processed
This is a processed version of BytedTsinghua-SIA/DAPO-Math-17k where we have:
Deduplicated the prompts
Reformatted the prompts and ground truth answers to be compatible with TRL's GRPO trainer
We have also derived pure English and Chinese subsets.
The full dataset processing logic can be found in create_dataset.py.
If you find this dataset useful in your work, please cite the original source with:
@misc{yu2025dapoopensourcellmreinforcement… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/DAPO-Math-17k-Processed.DAPO-Math-17k-Processed_filteredDAPO-Math-17K-cleanedQuestions and solutions for https://huggingface.co/datasets/BytedTsinghua-SIA/DAPO-Math-17k.
DAPO-MATH-17k-oss-reasoning
DAPO-MATH-17k-oss-reasoning
This dataset contains reasoning trajectories produced by gpt-oss-120b on BytedTsinghua-SIA/DAPO-Math-17k.
Under different reasoning efforts, we observe different token usage.
Effort Level
Avg Tokens
Low
1300
Medium
2936
High
8419
Keywords appearance frequency indicates reasoning efforts of the LLM.
Keyword
Low
Medium
High
All (L+M+H)
wait
40.5%
69.0%
87.3%
65.6%
double check
0.1%
2.4%
0.8%
1.9%
check
57.9%
87.3%… See the full description on the dataset page: https://huggingface.co/datasets/thuzhizhi/DAPO-MATH-17k-oss-reasoning.DAPO-Math-17k-dedupThis dataset is a simple deduplication of https://huggingface.co/datasets/BytedTsinghua-SIA/DAPO-Math-17k.
The deduplication is performed via this SQL: https://huggingface.co/datasets/BytedTsinghua-SIA/DAPO-Math-17k/sql-console/bES77R0.
dapo-math-17k-verl
DAPO-Math-17K VERL
📊 Dataset Summary
This dataset contains 17,147 mathematical reasoning problems in VERL format, processed from haizhongzheng/DAPO-Math-17K-cleaned.
Key Features:
17,147 high-quality math problems
Converted to VERL format for reward modeling
Verified ground truth answers
Ready for reinforcement learning training
🔗 Source Dataset
Original Repository
Repository: haizhongzheng/DAPO-Math-17K-cleaned
License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/dapo-math-17k-verl.DAPO-Math-17k-Processed_contaminatedDAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-non-thinking-dedup
DAPO Math Qwen3-235B non-thinking, deduplicated
This dataset is derived from
Yang-Zhou/DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill, specifically
dapo_distill_boxed_non_thinking.json.
It retains the original LLaMA-Factory-compatible instruction, input, and
output columns. Duplicate rows are identified by collapsing consecutive
whitespace in instruction, trimming leading/trailing whitespace, and hashing
the normalized instruction. The first row in each duplicate… See the full description on the dataset page: https://huggingface.co/datasets/YYYYYYibo/DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-non-thinking-dedup.dapo-math-17k-difficulty-qwen3-1.7b-base-k16
DAPO-Math-17k difficulty under Qwen3-1.7B-Base (K=16)
For each of the 17,398 problems in the DAPO-Math-17k train set, how many of
K=16 samples from the untrained base model are correct.
The headline: 57.27% of problems are solved 0 out of 16 times, and not one
problem is solved 16 out of 16. Difficulty here is entirely one-sided.
Why count per problem instead of reporting mean accuracy
In group-relative RL (GRPO and its relatives), a prompt group whose K responses… See the full description on the dataset page: https://huggingface.co/datasets/RyanYr/dapo-math-17k-difficulty-qwen3-1.7b-base-k16.DAPO-Math-17k-Processed-ScoredDAPO-Math-17k_gemini3-flash_qwen3-4B-Base-n4_sft_mmluprodapo-math-17k-qwen3-1.7b-base-n8
DAPO-Math-17k sampled with Qwen3-1.7B-Base, n=8
17398 problems from the RL training set, each sampled 8 times and scored with
the reward function the RL runs themselves used.
The point of this dataset is to be comparable with what the RL runs actually saw, so
every sampling knob is taken from the live training config or from the default that
config falls through to. Two of them are not in the config file at all and would be
wrong if guessed: top_k = -1 and min_tokens = 1.… See the full description on the dataset page: https://huggingface.co/datasets/RyanYr/dapo-math-17k-qwen3-1.7b-base-n8.DAPO-Math-17kDAPO-Math-17k-MATH-500DAPO-Math-17k with the MATH-500 test split converted to the same parquet schema and prompt format.
dapo-math-17kDAPO-Math-17k-easy6k-Qwen3-4B-k8DAPO-Math-17k-mixed-Qwen3-4B-k8dapo-math-17k-qwen3
DAPO-Math-17k English Qwen3 Answers
This dataset is derived from the English subset of open-r1/DAPO-Math-17k-Processed.
Added answer columns:
qwen3_1b7_answer
qwen3_4b_answer
qwen3_8b_answer
qwen3_14b_answer
qwen3_32b_answer
DAPO_Math_17k-gemini3-flashDAPO-Math-17kDAPO-Math-17k_simple_jjeDAPO-Math-17k_gemini3-flash_qwen3-4B-Base-n4_1e-5_sft_mathevaldapo-math-17k-Qwen3-1.7B-2kDAPO-Math-17k-Processed_filtered_olmo_completionsDAPO-Math-17k_jjedapo-math-17k-qwen3-cleanDAPO-Math-17k-ProcessedDAPO-Math-17k-split
