CoolFace
Modelpublic

bhaveshsoni0023/llama-3.2-1b-grpo-gsm8k

sourceHugging Facellama3.2updated 1mo agoView on Hugging Face
0likes143downloads
Model Card

Llama-3.2-1B-Instruct GRPO GSM8K

Llama-3.2-1B-Instruct trained with GRPO using NVIDIA NeMo RL v0.7.0 on a single H100.

Results

ModelGSM8K test (pass@1)Correct
meta-llama/Llama-3.2-1B-Instruct44.50%587 / 1319
This model51.48%679 / 1319

+7.0 points, 1.16x relative. Both models evaluated identically: full 1319-problem GSM8K test split, greedy decoding (temperature 0.0), same chain-of-thought prompt template, same eval harness.

Training

  • —GRPO, best checkpoint at step 190 of 200
  • —16 prompts x 8 generations per step (3,040 samples consumed)
  • —Reward: binary exact-match on the answer inside boxed tags
  • —DTensor v2 trainer + vLLM generation, colocated on one GPU
  • —1x NVIDIA H100 80GB, ~3 hours
  • —lr 1e-6 AdamW, KL penalty 0.01, clip 0.2/0.2, sequence length 1024
  • —Rollout temperature 1.0 (for exploration diversity)

No human-written solutions were used. The model learned from a programmatic grader checking whether its final answer matched ground truth.

A note on the size of the gain

This gain is smaller than GRPO runs starting from a base model, and that is expected. Llama-3.2-1B-Instruct has already been through SFT and RLHF, so it already emitted well-formatted step-by-step answers before training. A base model gains a large amount simply by learning the required output format; that portion was already present here.

For comparison, the same recipe applied to Qwen2.5-1.5B (a base model) gave 35.03% to 73.84%. Much of that larger jump was format compliance rather than reasoning. The +7.0 points here is closer to a measure of pure reasoning improvement.

In-training validation reported 60.94% on a 64-sample subset. The full test set gives 51.48%. The difference is checkpoint-selection bias: step 190 was chosen because it scored highest on those particular 64 questions. The 51.48% figure is the honest one.

Prompt format

The model was trained with this template and expects it at inference. A bare question without the wrapper will degrade results.

Think step-by-step to solve the following problem. Output your answer inside of \boxed{} tags.: {question}

Let's think step-by-step

Usage

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

m = "bhaveshsoni0023/llama-3.2-1b-grpo-gsm8k"
tok = AutoTokenizer.from_pretrained(m)
model = AutoModelForCausalLM.from_pretrained(m, dtype=torch.bfloat16, device_map="cuda")

q = "Natalia sold clips to 48 friends in April, and half as many in May. How many altogether?"
prompt = (f"Think step-by-step to solve the following problem. "
          f"Output your answer inside of \\boxed{{}} tags.:\n{q}\n\nLet's think step-by-step")

ids = tok(prompt, return_tensors="pt").to("cuda")
out = model.generate(**ids, max_new_tokens=512, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))

Use temperature 0.0 for evaluation, or ~0.6 with top_p 0.95 for production. Do not serve at temperature 1.0 — that was a training setting for exploration, not inference.

Limitations

Trained and evaluated only on grade-school arithmetic word problems. Performance on competition maths, formal proofs, or other domains has not been measured. The model is sensitive to prompt format.