dapomath17k
DAPO-Math-17kdapo-math-17kDAPO-Math-17k-Processed
Dataset Card for DAPO-Math-17k-Processed
This is a processed version of BytedTsinghua-SIA/DAPO-Math-17k where we have:
Deduplicated the prompts
Reformatted the prompts and ground truth answers to be compatible with TRL's GRPO trainer
We have also derived pure English and Chinese subsets.
The full dataset processing logic can be found in create_dataset.py.
If you find this dataset useful in your work, please cite the original source with:
@misc{yu2025dapoopensourcellmreinforcement… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/DAPO-Math-17k-Processed.DAPO-Math-17k-Processed_filteredDAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill
DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill
A high-quality Chain-of-Thought (CoT) dataset generated using Qwen/Qwen3-235B-A22B-Thinking-2507 with rejection sampling on BytedTsinghua-SIA/DAPO-Math-17k. This dataset is ideal for SFT distillation training to improve mathematical reasoning capabilities of models.
The dataset format is compatible with LLaMA-Factory for efficient SFT training.
Files
dapo_distill_boxed.json: Single sampling subset (15,129… See the full description on the dataset page: https://huggingface.co/datasets/Yang-Zhou/DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill.DAPO-Math-17K-cleanedQuestions and solutions for https://huggingface.co/datasets/BytedTsinghua-SIA/DAPO-Math-17k.
