ggbetz/qwen3-4b-think-s1-full-sft
0355
modrill/qwen3-4b-think-s1-full-sft
Full-parameter supervised fine-tuning (SFT) of Qwen/Qwen3-4B-Base on the think_s1 curriculum stage (easy + medium code reasoning data). This checkpoint is Stage1 of the think curriculum (checkpoint-1094).
Related models
- Think baseline full SFT: modrill/qwen3-4b-think-baseline-full-sft
- Nothink baseline full SFT: modrill/qwen3-4b-nothink-baseline-full-sft
Training summary
Eval (EvalScope, release_latest / AIME)
Usage
HuggingFace Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "modrill/qwen3-4b-think-s1-full-sft"
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id, trust_remote_code=True, torch_dtype="auto", device_map="auto"
)vLLM
python -m vllm.entrypoints.openai.api_server \
--model modrill/qwen3-4b-think-s1-full-sft \
--served-model-name think-s1 \
--max-model-len 32768 \
--port 8801Inference tips
- Use Qwen3 chat template with thinking enabled
- Recommended eval
max_tokens: 16384 (matches training cutoff) - Sampling: temperature=0.6, topp=0.95, topk=20
License
Apache 2.0, consistent with the Qwen3 base model license.
