RationalPursuit/Qwen3-4B-Stratos-GPT-OSS-120B-SFT
0
Model Card
Experimental research artifact. Base: `Qwen/Qwen3-4B-Instruct-2507`. SFT on `bespokelabs/Bespoke-Stratos-17k` gpt-oss-120b traces (reasoning effort medium, temperature 1.0, 14,768 paired rows).
Eval (Stratos held-out, 508 prompts): non-termination 32/508 = 0.063; complete <think> 476/508.
This is the third lineage in a set trained on the same prompts under the same recipe, varying only the teacher.
Out-of-Scope Use
Research artifact only — not intended for production use.
Framework versions
- PEFT 0.20.0
