CoolFace
Modelpublic

code-critic-model/Qwen3-4B-Critic-SFT-DPO

sourceHugging Faceapache-2.0updated 12d agoView on Hugging Face
0likes799downloads
Model Card

Qwen3-4B-Critic-SFT-DPO

The SFT + DPO critic from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It starts from Qwen3-4B-Critic-SFT and is further trained with direct preference optimization on pairs of the SFT critic's own critiques. It is the strongest 4B critic in the paper and the one reported in the + SFT + DPO rows of Table 1.

A critic sits next to a frozen coding agent. Every k agent steps it reads the trajectory so far and returns a short structured critique: which error categories it detects, the evidence, a recovery action, the task status, and one line of overall guidance. It steers the agent; it does not write the patch.

All released models and datasets are listed on the organization page. Code and configs are in the critic-training repository.

Where it appears in the paper

Paper locationRow label
Table 1, every agent blockQwen3-4B + SFT + DPO
Table 8, significance teststhe SFT + DPO critic
Section 2.4 and Figure 3the DPO training pipeline

Original checkpoint name: Qwen3-4B-SFT-DPO-4B-1409i-beta0.15-sft0.3-lr1e-6-bs32-ep3-step-80. The old name still redirects here.

DPO data

Preference pairs were built as described in Section 2.4 of the paper. The coding agent runs on training tasks; every k steps the SFT critic samples N=10 critiques for the current trajectory prefix; Claude Opus 4.6 acts as judge and picks the best and the worst critique by the correctness and clarity of their overall guidance. The best becomes chosen, the worst rejected. This checkpoint was trained on 1,409 such pairs, split 90/10 into train and evaluation.

The 1,409 pairs are not part of this release yet. The dataset code-critic-model/PRM_1541i is an earlier pair set built with the same procedure; it was used for development runs and is not the set behind this checkpoint.

Training setup

DPO with TRL, initialized from Qwen3-4B-Critic-SFT.

SettingValue
Initializationcode-critic-model/Qwen3-4B-Critic-SFT
ObjectiveDPO with an added SFT term on the chosen response, weight 0.3
beta0.15
Learning rate1e-6
Effective batch size32
Schedule3 epochs planned (120 steps); this checkpoint is step 80, the end of epoch 2
Precisionbf16

At step 80 the held-out preference accuracy was 0.68. The step-120 checkpoint is kept in the organization for reference but was not selected for the paper.

Results

Resolve rate on SWE-bench Verified (500 instances), from Table 1 of the paper, best of k=5 and k=10 per configuration.

Coding agentNo critic+ Qwen3-4B-Critic-SFT+ Qwen3-4B-Critic-SFT-DPO
Qwen3-32B8.811.414.4
Qwen3-Next-80B-A3B20.024.226.2
GPT-OSS-20B3.09.814.8
GLM-4.7-Flash-30B-A3B21.635.235.8
GPT-OSS-120B (medium reasoning)20.431.234.8
o3-mini19.027.628.2

DPO improves over SFT for all six agents. On Qwen3-32B, GPT-OSS-20B, and GPT-OSS-120B the 4B DPO critic also beats the 8B SFT critic.

How to use

Serve with vLLM in bf16 and run an agent through the repository's mini-swe-agent fork, which inserts a critique every k steps. The --prm name goes to LiteLLM, which needs a matching entry in mini-swe-agent/configs/litellm_model_registry.json to price the calls; copy one of the existing critic blocks to a new key Qwen3-4B-Critic-SFT-DPO. Without an entry the critic call fails and the agent runs without critiques.

bash
vllm serve code-critic-model/Qwen3-4B-Critic-SFT-DPO \
    --served-model-name Qwen3-4B-Critic-SFT-DPO \
    --dtype bfloat16 --max-model-len 65536 --port 8071

bash scripts/run_critic_max150.sh prm_issue_res_instructions_step_aware 5 0 qwen3-80b \
    --prm Qwen3-4B-Critic-SFT-DPO --prm-node <vllm-host>:8071 --slice :500 \
    --prefix-dir <path to the matching no-critic run>

To call the critic directly, take any record from critic-sft-cwm-qwen, drop its final teacher critique, and generate. A complete snippet is on the Qwen3-8B-Critic-SFT card; only the repo name changes.

Citation

bibtex
@misc{gandhi2026steerdontsolvetraining,
  title={Steer, Don't Solve: Training Small Critic Models for Large Code Agents},
  author={Shubham Gandhi and Yiqing Xie and Atharva Naik and Ruichen Zhu and Carolyn Rose},
  year={2026},
  eprint={2606.21811},
  archivePrefix={arXiv},
  primaryClass={cs.SE},
  url={https://arxiv.org/abs/2606.21811}
}