CoolFace
Modelpublic

CharlieLLL/Qwen3.5-9B-coding-solo350-raw-ckpt149-orch-eval150-20260920

sourceHugging Faceapache-2.0updated 7d agoView on Hugging Face
0likes10downloads
Model Card

Qwen3.5-9B coding worker: solo350-raw, ckpt149

Exact Hugging Face export evaluated in the 2026-09-20 SWE-bench Verified held-out eval150 campaign. Checkpoint is final (150 optimizer updates). Training mode: Solo RL. Evaluation mode: orch, MiniMax-M2.7 orchestrator plus this 9B worker; worker thinking disabled.

Two new independent full150 repetitions use 32 concurrent episodes and 10GiB sandboxes. All seven models use the same held-out tasks and frozen regression-gated evaluator. Previous 12-concurrency measurements are historical, separately labeled. No claim of stable improvement follows from a single score.

Full results, task latency, completion T25/T50/T75/T90/T100, Docker breakdown, role-separated token and prefix-cache accounting, traces, and matched baselines: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920

The model files are inference weights; optimizer/RNG training-resume state is not part of this inference export. ORIGINALCHECKPOINT.json records provenance. MODELSHA256.json lists every published model/config file. eval_chat_template.jinja is the exact evaluation template; the native chat_template.jinja is retained separately. Use the evaluation template and disable thinking to reproduce this campaign.

Origin revision: c202236235762e1c871ad0ccb60c8ee5ba337b9a. Coordinator revision: d494266a4affc0d2995ba1fa35c8481cbd84294b. Checkpoint numbers are zero-based local saved iteration numbers and must not be confused with the OPD109 initialization label.