bhaveshsoni0023/qwen2.5-1.5b-grpo-gsm8k
0116
Qwen2.5-1.5B GRPO GSM8K
Qwen2.5-1.5B trained with GRPO using NVIDIA NeMo RL v0.7.0.
Results
+38.8 points, 2.1x relative. Both models evaluated identically on the full 1319-problem GSM8K test split, greedy decoding, same prompt template.
Training
- GRPO, 130 steps, 16 prompts x 8 generations (2,080 samples)
- Reward: binary exact-match on the answer inside boxed tags
- DTensor v2 trainer + vLLM generation, colocated
- 1x H100 80GB, ~3 hours
- lr 1e-6 AdamW, KL penalty 0.01, clip 0.2/0.2, seq len 1024
No human-written solutions. The model learned from a programmatic grader.
Prompt format
Trained with this template and expects it at inference. A bare question without the wrapper degrades results substantially.
Think step-by-step to solve the following problem. Output your answer inside of \boxed{} tags.: {question}
Let's think step-by-step
