datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3-4b-teacher-rollouts-76k-nonthinking
Qwen3-4B Teacher Rollouts 76K Non-Thinking
This dataset contains 76,800 fixed teacher trajectories generated for a
prompt-aligned reproduction study of on-policy distillation with Qwen3-1.7B.
It is an independent research artifact, not an official release from the model
or paper authors.
Models and generation
Teacher: Qwen/Qwen3-4B-Instruct-2507
Tokenizer/chat template: Qwen/Qwen3-1.7B
Mode: non-thinking (enable_thinking=False)
Temperature: 0.7
Top-p: 1.0
Top-k:… See the full description on the dataset page: https://huggingface.co/datasets/YangyiH/qwen3-4b-teacher-rollouts-76k-nonthinking.qwen3-4b-nonthinking-skywork-rollouts-verified
Qwen3-4B Nonthinking Skywork Math Rollouts
Rollouts generated by Qwen/Qwen3-4B in nonthinking mode. Answers are verified with
math-verify==0.9.0 after extracting the final Answer: or \boxed{} value. Bare LaTeX and text answers are retried inside a math environment before verify.
Rows: 215,040
Prompt groups: 26,880
Samples per prompt: 8
Correct rows: 92,497
Overall success rate: 0.430139
Each row contains binary reward, the success fraction across all samples sharing the
same… See the full description on the dataset page: https://huggingface.co/datasets/YYYYYYibo/qwen3-4b-nonthinking-skywork-rollouts-verified.
