bhaveshsoni0023/llama-3.2-1b-grpo-gsm8k
Llama-3.2-1B-Instruct GRPO GSM8K
Llama-3.2-1B-Instruct trained with GRPO using NVIDIA NeMo RL v0.7.0 on a single H100.
Results
+7.0 points, 1.16x relative. Both models evaluated identically: full 1319-problem GSM8K test split, greedy decoding (temperature 0.0), same chain-of-thought prompt template, same eval harness.
Training
- GRPO, best checkpoint at step 190 of 200
- 16 prompts x 8 generations per step (3,040 samples consumed)
- Reward: binary exact-match on the answer inside boxed tags
- DTensor v2 trainer + vLLM generation, colocated on one GPU
- 1x NVIDIA H100 80GB, ~3 hours
- lr 1e-6 AdamW, KL penalty 0.01, clip 0.2/0.2, sequence length 1024
- Rollout temperature 1.0 (for exploration diversity)
No human-written solutions were used. The model learned from a programmatic grader checking whether its final answer matched ground truth.
A note on the size of the gain
This gain is smaller than GRPO runs starting from a base model, and that is expected. Llama-3.2-1B-Instruct has already been through SFT and RLHF, so it already emitted well-formatted step-by-step answers before training. A base model gains a large amount simply by learning the required output format; that portion was already present here.
For comparison, the same recipe applied to Qwen2.5-1.5B (a base model) gave 35.03% to 73.84%. Much of that larger jump was format compliance rather than reasoning. The +7.0 points here is closer to a measure of pure reasoning improvement.
In-training validation reported 60.94% on a 64-sample subset. The full test set gives 51.48%. The difference is checkpoint-selection bias: step 190 was chosen because it scored highest on those particular 64 questions. The 51.48% figure is the honest one.
Prompt format
The model was trained with this template and expects it at inference. A bare question without the wrapper will degrade results.
Think step-by-step to solve the following problem. Output your answer inside of \boxed{} tags.: {question}
Let's think step-by-step
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
m = "bhaveshsoni0023/llama-3.2-1b-grpo-gsm8k"
tok = AutoTokenizer.from_pretrained(m)
model = AutoModelForCausalLM.from_pretrained(m, dtype=torch.bfloat16, device_map="cuda")
q = "Natalia sold clips to 48 friends in April, and half as many in May. How many altogether?"
prompt = (f"Think step-by-step to solve the following problem. "
f"Output your answer inside of \\boxed{{}} tags.:\n{q}\n\nLet's think step-by-step")
ids = tok(prompt, return_tensors="pt").to("cuda")
out = model.generate(**ids, max_new_tokens=512, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))Use temperature 0.0 for evaluation, or ~0.6 with top_p 0.95 for production. Do not serve at temperature 1.0 — that was a training setting for exploration, not inference.
Limitations
Trained and evaluated only on grade-school arithmetic word problems. Performance on competition maths, formal proofs, or other domains has not been measured. The model is sensitive to prompt format.
