CharlieLLL/Qwen3.5-9B-coding-graded350-opd-ckpt149-appendonly-w64-20260920
Qwen3.5-9B coding worker: graded350-opd, ckpt149
Exact Hugging Face export evaluated in the 2026-09-20 SWE-bench Verified held-out eval150 campaign. Checkpoint is final (150 optimizer updates). Training mode: Graded RL. Evaluation mode: orch, MiniMax-M2.7 orchestrator plus this 9B worker; worker thinking disabled.
One independent full150 screening uses 64 concurrent episodes and 10GiB sandboxes. All four models use the same held-out tasks and append-only regression-gated evaluator. Prior 32-concurrency/original-harness results are separately labeled historical context, not matched baselines. No claim of stable improvement follows from a single score.
Full results, task latency, completion T25/T50/T75/T90/T100, Docker breakdown, role-separated token and prefix-cache accounting, traces, and matched baselines: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-new-checkpoints-appendonly-w64-20260920
The model files are inference weights; optimizer/RNG training-resume state is not part of this inference export. ORIGINALCHECKPOINT.json records provenance. MODELSHA256.json lists every published model/config file. eval_chat_template.jinja is the exact evaluation template; the native chat_template.jinja is retained separately. Use the evaluation template and disable thinking to reproduce this campaign.
Export donor revision: 679d1265badc03391b17126b79096ae3b6ffb20c. Training initialization is recorded separately in ORIGINAL_CHECKPOINT.json; Solo350 local63 is not the raw-model baseline. Coordinator revision: d494266a4affc0d2995ba1fa35c8481cbd84294b. Checkpoint numbers are zero-based local saved iteration numbers and must not be confused with the OPD109 initialization label.
