mariklolik228/sus-qwen2.5-1.5b-grpo-lora
025
SuS-GRPO LoRA — Qwen2.5-1.5B-Instruct
LoRA adapter fine-tuned with SuS (Strategy-aware Surprise) + GRPO for math reasoning on GSM8K.
Method
SuS adds a semantic novelty bonus to correct responses in mixed-correctness batches, encouraging diverse problem-solving strategies. Zero-variance batches (all-correct or all-incorrect) are left untouched — this prevents KL divergence blowup and length collapse.
- Base model: Qwen/Qwen2.5-1.5B-Instruct
- Encoder: all-MiniLM-L6-v2 (frozen, 22M params)
Training Configuration
Results on GSM8K
95% CI (Pass@1): [73.49, 77.05]
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct", torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(base, "mariklolik228/sus-qwen2.5-1.5b-grpo-lora")
tokenizer = AutoTokenizer.from_pretrained("mariklolik228/sus-qwen2.5-1.5b-grpo-lora")Links
- Paper: SuS: Strategy-aware Surprise for Intrinsic Exploration in GRPO (arXiv:2601.10349)
- Code: github.com/mariklolik/sus
- Training Logs: W&B
- Baseline: mariklolik228/grpo-baseline-qwen2.5-1.5b-lora
Citation
@article{kashirskiy2026sus,
title={SuS: Strategy-aware Surprise for Intrinsic Exploration in GRPO},
author={Kashirskiy, Mark and Makarov, Ilya},
journal={arXiv preprint arXiv:2601.10349},
year={2026}
}