datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DAPO-Math-17kDAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill
DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill
A high-quality Chain-of-Thought (CoT) dataset generated using Qwen/Qwen3-235B-A22B-Thinking-2507 with rejection sampling on BytedTsinghua-SIA/DAPO-Math-17k. This dataset is ideal for SFT distillation training to improve mathematical reasoning capabilities of models.
The dataset format is compatible with LLaMA-Factory for efficient SFT training.
Files
dapo_distill_boxed.json: Single sampling subset (15,129… See the full description on the dataset page: https://huggingface.co/datasets/Yang-Zhou/DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill.dapo-math-17k-verl
DAPO-Math-17K VERL
📊 Dataset Summary
This dataset contains 17,147 mathematical reasoning problems in VERL format, processed from haizhongzheng/DAPO-Math-17K-cleaned.
Key Features:
17,147 high-quality math problems
Converted to VERL format for reward modeling
Verified ground truth answers
Ready for reinforcement learning training
🔗 Source Dataset
Original Repository
Repository: haizhongzheng/DAPO-Math-17K-cleaned
License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/dapo-math-17k-verl.DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-non-thinking-dedup
DAPO Math Qwen3-235B non-thinking, deduplicated
This dataset is derived from
Yang-Zhou/DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill, specifically
dapo_distill_boxed_non_thinking.json.
It retains the original LLaMA-Factory-compatible instruction, input, and
output columns. Duplicate rows are identified by collapsing consecutive
whitespace in instruction, trimming leading/trailing whitespace, and hashing
the normalized instruction. The first row in each duplicate… See the full description on the dataset page: https://huggingface.co/datasets/YYYYYYibo/DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-non-thinking-dedup.dapo-math-17k-qwen3-1.7b-base-n8
DAPO-Math-17k sampled with Qwen3-1.7B-Base, n=8
17398 problems from the RL training set, each sampled 8 times and scored with
the reward function the RL runs themselves used.
The point of this dataset is to be comparable with what the RL runs actually saw, so
every sampling knob is taken from the live training config or from the default that
config falls through to. Two of them are not in the config file at all and would be
wrong if guessed: top_k = -1 and min_tokens = 1.… See the full description on the dataset page: https://huggingface.co/datasets/RyanYr/dapo-math-17k-qwen3-1.7b-base-n8.DAPO-Math-17kDAPO-Math-17k-MATH-500DAPO-Math-17k with the MATH-500 test split converted to the same parquet schema and prompt format.
DAPO-Math-17kDAPO-Math-17kDAPO-Math-17k
