lorenzocazzador/heavy-reviewer-simpo-qwen3-32b
012
Heavy Reviewer (multi-task SimPO) — Qwen3-32B LoRA
A single LoRA adapter on `Qwen/Qwen3-32B` that acts as the heavy (high-capacity) reviewer in a self-refining coding/math agent loop. One shared adapter drives two heads, generate-then-parse:
- Verdict — judge a candidate solution (accept / revise, with a critique);
- Test generation — propose discriminative tests that expose bugs in a candidate.
Trained with multi-task SimPO using an RPO anchor (the "v3" recipe), which holds logp(chosen) up and keeps both generators coherent — unlike likelihood-/length-collapsed variants. Net effect: +11.46 pp verdict scoring accuracy over the base model while both heads stay usable (verdict and test-gen spot-checks pass and are substantively correct).
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = "Qwen/Qwen3-32B"
repo = "lorenzocazzador/heavy-reviewer-simpo-qwen3-32b"
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, repo)
tokenizer = AutoTokenizer.from_pretrained(repo)Adapter produced for a master's thesis on self-refining coding/math agents.
