violetxi/qwen35-9b-equational-theory-sair-mix100m
Qwen3.5-9B equational-theory SAIR — 100M mixture
Final epoch 2 / step 1,846 checkpoint from completed training job 1020938. This is a complete BF16 sharded safetensors model in the native Qwen3_5ForConditionalGeneration layout, including tokenizer, chat template, configuration and processors. No adapter merge or checkpoint conversion is required.
Training
- Base:
Qwen/Qwen3.5-9B, pinned revisionc202236235762e1c871ad0ccb60c8ee5ba337b9a. - 99,999,924 effective supervised tokens per epoch, two epochs: 199,999,848 supervised-token exposures.
- 103,276 usable notes from the R0–R4 snapshot contribute 60,927,286 tokens per epoch; 42,013 curated R0–R3 note-conditioned trajectories contribute 39,072,638 tokens. Token mixture: 60.9273% notes / 39.0727% trajectories.
- This uses a newer note snapshot and a different ratio from the earlier approximately 70/30 1M–30M ladder, so the comparison changes more than dataset size.
- Full-model SFT with FSDP2 on eight GH200 GPUs; learning rate 5e-6, cosine schedule, 3% warmup, BF16 computation, activation checkpointing, global supervised-token mean cross-entropy, no KL.
- Padding-free packing with isolated examples; prompts are masked in trajectory loss.
- Final held-out token NLL: 0.2791766 (notes 0.3221392, trajectories 0.1792197).
- Trained FP32 language weights are converted to BF16; all 427 language tensors are mapped into the native layout. The remaining 348 vision/MTP tensors are inherited from the pinned base. Standard Transformers model loading passed with no missing, unexpected or mismatched keys.
- Online training log.
Data
- Published note snapshot, revision
fbc728524ae3a2c8cc6a8806830698d3878fac7c. - Published note-conditioned rollouts, revision
19e5d3ec82bf73b89feb436cfe54ff87f4c1c106. Theselected_for_100mfield identifies the 42,013 selected trajectories. The archive also contains other responses that were not used in this training run.
Evaluation
As of publication on September 24, 2026, no completed SAIR Stage 1 score is available for this checkpoint. The initial inference job failed during CUDA library linking before generating responses; a corrected evaluation is queued. No Stage 2 result is reported.
The requested Stage 1 protocol is 800 questions × four seeded samples, official rule-based binary-verdict grading, 24,576 output tokens including thinking, and 32,768 context tokens. Seeds 0–3; temperature 1.0, top-p 0.95, top-k 20, min-p 0, presence penalty 1.5, frequency penalty 0, repetition penalty 1.0.
Loading
Use a Transformers release with Qwen3.5 support.
from transformers import AutoTokenizer, Qwen3_5ForConditionalGeneration
repo = "violetxi/qwen35-9b-equational-theory-sair-mix100m"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
repo, dtype="auto", device_map="auto"
)