CoolFace
Modelpublic

RationalPursuit/Qwen3-4B-Stratos-GPT-OSS-120B-SFT

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes
Model Card

Model Card

Experimental research artifact. Base: `Qwen/Qwen3-4B-Instruct-2507`. SFT on `bespokelabs/Bespoke-Stratos-17k` gpt-oss-120b traces (reasoning effort medium, temperature 1.0, 14,768 paired rows).

Eval (Stratos held-out, 508 prompts): non-termination 32/508 = 0.063; complete <think> 476/508.

This is the third lineage in a set trained on the same prompts under the same recipe, varying only the teacher.

Out-of-Scope Use

Research artifact only — not intended for production use.

Framework versions

  • —PEFT 0.20.0