CoolFace
Modelpublic

code-critic-model/Qwen3-4B-SFT-DPO-4B-1409i-beta0.15-sft0.3-lr1e-6-bs32-ep3-step-120

sourceHugging Faceupdated 25d agoView on Hugging Face
0likes371downloads
Model Card
Not used in the paper. This is a development checkpoint kept so that existing links keep working. The critics reported in Steer, Don't Solve are listed on the organization page; the released DPO critic is Qwen3-4B-Critic-SFT-DPO.

Qwen3-4B-SFT-DPO-4B-1409i-beta0.15-sft0.3-lr1e-6-bs32-ep3-step-120

Final checkpoint (step 120, end of epoch 3) of the DPO run whose step-80 checkpoint was released as Qwen3-4B-Critic-SFT-DPO. Same initialization, data, and hyperparameters (beta 0.15, SFT weight 0.3, lr 1e-6, effective batch 32); only the number of steps differs. Held-out preference accuracy at step 120 was 0.65 versus 0.68 at step 80, and the step-80 checkpoint was selected for the paper.