Parallel-R1/Parallel-R1-Unseen_Step_200
060
1---2license: mit3datasets:4- Leo-Dai/dapo-math-17k_dedup5---6# 🧠 Parallel-R1-Unseen_Step_2007 8> **Mid-Training Checkpoint of Parallel-R1: Towards Parallel Thinking via Reinforcement Learning** 9> Stage: **After 200 RL steps via alternating rewards** — showing the adaptive parallel reasoning ability and serve as structure exploration stage.10 11This checkpoint aims to help you reproduce experimental results in Section 4.5: Extra Bonus: Parallel Thinking as a Mid-Training Exploration Strategy for RL Training.12 13 