AlexWortega/moe-200m-qwen3-step160000-5B-20260522-2256
MoE-200M (Qwen3) — step 160 000 / 100B-run intermediate checkpoint (5.24 B tokens)
Mid-training snapshot of the in-flight moe-200m-qwen3-100b pretrain run, captured at step 159 999 ≈ 5.24 B tokens consumed (5.2 % of the 100 B-token budget, +74 % more training than the step 70 000 / 2.4 B snapshot). Trained autonomously by the ml-intern Claude Code skill on 2× Tesla V100-SXM2 32 GB.
Source code, training log, full eval bundle: **AlexWortega/moe-200m-qwen3-100b-** (GitHub repo, in progress).
The model is not yet converged — the 100 B run continues. The final checkpoint will land at AlexWortega/moe-200m-qwen3-100b-* once the run completes.
Architecture
DeepSeekMoE-style MoE with a Qwen3 tokenizer:
The MoE dispatch uses a token-permuted, capacity-padded grouped-bmm kernel (moe_backend="grouped") — stacked expert weights of shape [E, d_ff, d] (gate, up) and [E, d, d_ff] (down). State-dict keys are flat blocks.{i}.ffn.{gate,up,down} rather than per-expert ModuleList entries; the legacy per-expert layout is also accepted (auto-stacked) by MoEModel.load_state_dict.
The router is a sigmoid + top-k pick with aux_coef=1e-3, z_coef=1e-3, and an additive bias controller (symmetric, with starved-expert boost, clamped to ±10). See model.py for the full SigmoidRouter + _moe_dispatch_grouped implementation.
Training state @ step 160 000
Zero-shot lm-evaluation-harness results
6 standard tasks, num_fewshot=0, batch_size=8, dtype=float16, Qwen3 tokenizer, single seed, no bootstrap. Δ is vs the step 90 000 / 3.0 B snapshot of the same run.
Average Δ over the 6 headline tasks vs the 3 B snapshot: +1.9 pt (3 of 6 positive). Boolq alone added +12.9 pt; piqa / arc_easy / lambada slipped by 0.5–1.9 pt (within single-seed noise band on these set sizes).
The headline qualitative event in this 2.24 B-extra-tokens window is boolq crossing chance (44 → 57) — eval_loss is still dropping monotonically (3.55 → 3.46) and lambada perplexity is now finite (was overflowing fp16 on our 100M @ 21B run), but most multi-choice tasks haven't moved much yet.
See EVAL_5B.md in the run-dir / GitHub repo for the full honesty pass.
How to load
A self-contained, runnable test ships with the repo:
pip install transformers safetensors huggingface_hub torch
python load_test.pyload_test.py does the full reload: snapshot-downloads this repo, imports MoEModel from the bundled model.py, builds the config from config.json, loads model.safetensors strictly, runs one forward pass, and samples 60 tokens from "Once upon a time,". A successful run prints LOAD_TEST: PASS.
Python:
from huggingface_hub import snapshot_download
import importlib.util, json, torch
from safetensors.torch import load_file
from pathlib import Path
local = Path(snapshot_download("AlexWortega/moe-200m-qwen3-step160000-5B-20260522-2256"))
spec = importlib.util.spec_from_file_location("_mdl", local / "model.py")
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
cfg = json.loads((local / "config.json").read_text())
cfg = {k: v for k, v in cfg.items() if not k.startswith("_")}
config = mod.MoEModelConfig(**cfg)
config.router_noise_std = 0.0
config.use_liger_ce = False
config.use_chunked_ce = False
model = mod.MoEModel(config).to("cuda").to(torch.float32)
sd = load_file(local / "model.safetensors", device="cuda")
missing, unexpected = model.load_state_dict(sd, strict=False)
assert not unexpected
model.eval()Reproducibility
model.py is the verbatim training-time file. config.json is generated from the checkpoint's cfg dict (asdict on the MoEModelConfig dataclass) with _model_class, _ckpt_step, and _tokens_seen metadata appended. The Qwen3 tokenizer is the unmodified Qwen/Qwen3-0.6B-Base tokenizer.
Caveats
- Mid-training snapshot — outputs are coherent at the sentence level but not yet competitive on multi-choice benchmarks vs reference 124–500 M models.
- Sigmoid router + grouped-bmm dispatch is the only supported backend at this scale; the legacy per-expert bmm backend would fall over on V100 throughput.
router_noise_std/use_liger_ce/use_chunked_ceare forced false inconfig.jsonfor inference; the training-time AMP / chunked-CE / Liger fused-CE paths are documented in the bundledmodel.pybut are not exercised byload_test.py.
