CoolFace
Modelpublic

RockToken/qwen3_30b_a3b_to_4b_onpolicy_science_5k

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes11downloads
Model Card

Qwen3-4B distilled from Qwen3-30B-A3B — On-policy 5k (science)

On-policy KD for a Qwen3-4B student toward the Qwen3-30B-A3B teacher, on 5,000 science prompts from OpenThoughts-3.

Domain-transfer companion to the math-domain checkpoints in this organisation: same student starting point and recipe as the onpolicy_5k_src* chain, but the on-policy prompts come from the science slice instead of math. Training exposure of this checkpoint:

  1. 1.Off-policy KD on 20k math prompts → `qwen3_30b_a3b_to_4b_offpolicy_20k`
  2. 2.On-policy KD (this run) on 5k science prompts → this checkpoint

Models

RoleModel
StudentRockToken/qwen330ba3bto4boffpolicy20k
TeacherQwen/Qwen3-30B-A3B-Instruct-2507 (MoE)

enable_thinking=False throughout.

Training data

  • —Source: open-thoughts/OpenThoughts3-1.2M, domain == "science" slice
  • —5,000 single-user-turn prompts; 69 exceeding prompt_max_len=1280 were filtered, leaving 4,931
  • —Only the prompts are used; on-policy KD never reads the dataset's reference answers

Training setup

Framework: KDFlow — FSDP2 + SGLang rollout, Ray-orchestrated GPU co-location with sleep/wakeup. Same pipeline revision as the math-domain onpolicy_5k_src* runs.

Hardware: 1× node, 4× H100 (94 GB), 24 h 02 min wall-clock (~96 GPU-hours).

Key hyperparameters

GroupValue
Backendfsdp2, bf16, gradient ckpt on
Epochs1 (2,465 rollout iterations)
Train batch4 (micro 1)
Learning rate2e-6, cosine, warmup 5%
KD ratio1.0
KD lossreverse KL (rkl)
KD algorithmvanilla_kd
Rollout engineSGLang, TP=2, 1 engine
Rollout batch2 prompts × 4 samples/prompt
generate_max_len8000
prompt_max_len1280 (total max_len 9536)
Samplingtemperature 1.0, top-p 1.0
TeacherTP=4, sleep/wakeup enabled

prompt_max_len is 1280 instead of the math runs' 800: science prompts run longer (p99 ≈ 1,182 tokens).

Note on the LR schedule: as in the earlier runs of this series, the scheduler's horizon is sized to 4× the actual number of updates, so the cosine decay is only partially traversed — the LR moves from 2.00e-6 to 1.79e-6 over the run rather than reaching zero. Kept as is for comparability within the series.

Training dynamics

Means over 500-step blocks (per-step values are noisy at batch 4):

stepsloss (reverse KL)mean gen length
1–5000.2733,250
501–10000.2563,417
1001–15000.2533,456
1501–20000.2483,390
2001–24650.2423,321

The loss starts near 0.27 rather than the ~3 seen when distilling from a raw instruct model — the student has already been through off-policy KD against the same teacher, so on-policy training starts close and grinds out the remaining distribution gap.

Generation lengths: p50 = 3,267, p95 = 6,456, max = 8,000 against the 8,000 cap. A small tail of responses hits the cap; the bulk is untruncated.

Intended use

Research on distillation dynamics, in particular whether on-policy KD gains transfer across domains (math-trained off-policy base → science on-policy round). Domain: science (OpenThoughts-3 science split).

Limitations

  • —The on-policy round is science-only; not tuned for chat or safety.
  • —enable_thinking=False — this student does not emit thinking traces.
  • —A small fraction of rollouts was truncated at 8,000 tokens during training.
  • —Single checkpoint at the end of one epoch; no intermediate checkpoints were kept.