CoolFace
Modelpublic

lorenzocazzador/heavy-reviewer-simpo-qwen3-32b

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes12downloads
Model Card

Heavy Reviewer (multi-task SimPO) — Qwen3-32B LoRA

A single LoRA adapter on `Qwen/Qwen3-32B` that acts as the heavy (high-capacity) reviewer in a self-refining coding/math agent loop. One shared adapter drives two heads, generate-then-parse:

  • —Verdict — judge a candidate solution (accept / revise, with a critique);
  • —Test generation — propose discriminative tests that expose bugs in a candidate.

Trained with multi-task SimPO using an RPO anchor (the "v3" recipe), which holds logp(chosen) up and keeps both generators coherent — unlike likelihood-/length-collapsed variants. Net effect: +11.46 pp verdict scoring accuracy over the base model while both heads stay usable (verdict and test-gen spot-checks pass and are substantively correct).

Base modelQwen/Qwen3-32B
Methodmulti-task SimPO + RPO anchor (LoRA, r = 32, α = 64)
Headsverdict + discriminative test generation (single shared adapter)

Usage

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "Qwen/Qwen3-32B"
repo = "lorenzocazzador/heavy-reviewer-simpo-qwen3-32b"

model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, repo)
tokenizer = AutoTokenizer.from_pretrained(repo)

Adapter produced for a master's thesis on self-refining coding/math agents.