eewer/Qwen3-4B-Thinking-Preservation-terminus2-sft
Qwen3-4B-Thinking-Preservation-terminus2-sft
eewer/Qwen3-4B-Thinking-Preservation supervised-fine-tuned for terminal-agent use in the terminus-2 format (native JSON-in-text actions; thinking preserved across multi-turn history). This is the final checkpoint of the default run (globalstep 2790; ~1 epoch over the shuffled skill-based-medium terminus-2 SFT mix; AdamW, constant LR 5e-6 after a short warmup).
Why this checkpoint
It is the best checkpoint by our reliable terminal-bench eval (terminus-2 harness, 6 informative live tasks, n=15 trials/task, temp 0.6 / top_p 0.95, 8192 out tokens, 40 turns):
The raw final checkpoint beats every individual checkpoint and both stochastic-weight- averaging (SWA) merges — checkpoint merging gave no gain for this constant-LR run. Use it as a drop-in base for downstream training/eval.
