CoolFace
Modelpublic

eewer/Qwen3-4B-Thinking-Preservation-terminus2-sft

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes5downloads
Model Card

Qwen3-4B-Thinking-Preservation-terminus2-sft

eewer/Qwen3-4B-Thinking-Preservation supervised-fine-tuned for terminal-agent use in the terminus-2 format (native JSON-in-text actions; thinking preserved across multi-turn history). This is the final checkpoint of the default run (globalstep 2790; ~1 epoch over the shuffled skill-based-medium terminus-2 SFT mix; AdamW, constant LR 5e-6 after a short warmup).

Why this checkpoint

It is the best checkpoint by our reliable terminal-bench eval (terminus-2 harness, 6 informative live tasks, n=15 trials/task, temp 0.6 / top_p 0.95, 8192 out tokens, 40 turns):

checkpoint6-task pass rate
this (default-final, s2790)45.6%
default-s999 / s1499 / s199941.1% / 40.0% / 37.0%
SWA merges (full-tail / last-6)38.9% / 37.8%
diverse run (best / latest)35.6% / 26.6%

The raw final checkpoint beats every individual checkpoint and both stochastic-weight- averaging (SWA) merges — checkpoint merging gave no gain for this constant-LR run. Use it as a drop-in base for downstream training/eval.