CoolFace
Modelpublic

AlexWortega/moe-200m-qwen3-step160000-5B-20260522-2256

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes19downloads
Model Card

MoE-200M (Qwen3) — step 160 000 / 100B-run intermediate checkpoint (5.24 B tokens)

Mid-training snapshot of the in-flight moe-200m-qwen3-100b pretrain run, captured at step 159 999 ≈ 5.24 B tokens consumed (5.2 % of the 100 B-token budget, +74 % more training than the step 70 000 / 2.4 B snapshot). Trained autonomously by the ml-intern Claude Code skill on 2× Tesla V100-SXM2 32 GB.

Source code, training log, full eval bundle: **AlexWortega/moe-200m-qwen3-100b-** (GitHub repo, in progress).

The model is not yet converged — the 100 B run continues. The final checkpoint will land at AlexWortega/moe-200m-qwen3-100b-* once the run completes.

Architecture

DeepSeekMoE-style MoE with a Qwen3 tokenizer:

total params616.5 M
active params203.6 M / token
vocab151 936 (Qwen3)
d_model640
n_layers16 (layer 0 dense, layers 1–15 MoE)
attentionGQA — 10 Q heads / 2 KV heads, head_dim 64, partial RoPE (32 dims)
experts16 routed + 1 shared, top-2 sigmoid router
d_ff1024 (per expert)
tied embed/lm_headyes
µP base_d512
precision (training)fp16 AMP, Muon + AdamW, WSD schedule

The MoE dispatch uses a token-permuted, capacity-padded grouped-bmm kernel (moe_backend="grouped") — stacked expert weights of shape [E, d_ff, d] (gate, up) and [E, d, d_ff] (down). State-dict keys are flat blocks.{i}.ffn.{gate,up,down} rather than per-expert ModuleList entries; the legacy per-expert layout is also accepted (auto-stacked) by MoEModel.load_state_dict.

The router is a sigmoid + top-k pick with aux_coef=1e-3, z_coef=1e-3, and an additive bias controller (symmetric, with starved-expert boost, clamped to ±10). See model.py for the full SigmoidRouter + _moe_dispatch_grouped implementation.

Training state @ step 160 000

tokens_seen5 242 880 000 (≈ 5.2 % of 100 B target)
step159 999
train lm_loss (200-step window)≈ 3.47
eval loss (held-out)3.458
router CV≈ 0.64
router entropy≈ 3.61 bits
throughput≈ 26.5 k tokens/s on 2× V100
hardware2× Tesla V100-SXM2 32 GB (GPUs 2, 3)

Zero-shot lm-evaluation-harness results

6 standard tasks, num_fewshot=0, batch_size=8, dtype=float16, Qwen3 tokenizer, single seed, no bootstrap. Δ is vs the step 90 000 / 3.0 B snapshot of the same run.

taskmetricrandomgpt2-124Mpythia-160mour-100M@21BLFM2-350Mqwen2.5-0.5Bour-200M@3B**our-200M@5.2B (this)**Δ vs 3B
boolqacc50.048.755.258.164.262.544.156.9+12.9
hellaswagacc25.028.928.431.738.440.628.729.7+1.0
hellaswagacc_norm25.031.230.335.949.052.231.132.2+1.1
piqaacc_norm50.062.561.464.269.569.957.657.1−0.5
winograndeacc50.052.451.050.055.756.549.450.2+0.8
arc_easyacc25.043.643.854.969.464.544.742.8−1.9
arc_easyacc_norm25.039.639.948.966.258.642.242.1~0
lambada_openaiacc0.032.232.723.540.252.517.816.6−1.2
lambada_openaippl ↓—40.138.1overflow27.410.61064.3984.6−7.5 %

Average Δ over the 6 headline tasks vs the 3 B snapshot: +1.9 pt (3 of 6 positive). Boolq alone added +12.9 pt; piqa / arc_easy / lambada slipped by 0.5–1.9 pt (within single-seed noise band on these set sizes).

The headline qualitative event in this 2.24 B-extra-tokens window is boolq crossing chance (44 → 57) — eval_loss is still dropping monotonically (3.55 → 3.46) and lambada perplexity is now finite (was overflowing fp16 on our 100M @ 21B run), but most multi-choice tasks haven't moved much yet.

See EVAL_5B.md in the run-dir / GitHub repo for the full honesty pass.

How to load

A self-contained, runnable test ships with the repo:

bash
pip install transformers safetensors huggingface_hub torch
python load_test.py

load_test.py does the full reload: snapshot-downloads this repo, imports MoEModel from the bundled model.py, builds the config from config.json, loads model.safetensors strictly, runs one forward pass, and samples 60 tokens from "Once upon a time,". A successful run prints LOAD_TEST: PASS.

Python:

python
from huggingface_hub import snapshot_download
import importlib.util, json, torch
from safetensors.torch import load_file
from pathlib import Path

local = Path(snapshot_download("AlexWortega/moe-200m-qwen3-step160000-5B-20260522-2256"))
spec = importlib.util.spec_from_file_location("_mdl", local / "model.py")
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)

cfg = json.loads((local / "config.json").read_text())
cfg = {k: v for k, v in cfg.items() if not k.startswith("_")}
config = mod.MoEModelConfig(**cfg)
config.router_noise_std = 0.0
config.use_liger_ce = False
config.use_chunked_ce = False

model = mod.MoEModel(config).to("cuda").to(torch.float32)
sd = load_file(local / "model.safetensors", device="cuda")
missing, unexpected = model.load_state_dict(sd, strict=False)
assert not unexpected
model.eval()

Reproducibility

model.py is the verbatim training-time file. config.json is generated from the checkpoint's cfg dict (asdict on the MoEModelConfig dataclass) with _model_class, _ckpt_step, and _tokens_seen metadata appended. The Qwen3 tokenizer is the unmodified Qwen/Qwen3-0.6B-Base tokenizer.

Caveats

  • —Mid-training snapshot — outputs are coherent at the sentence level but not yet competitive on multi-choice benchmarks vs reference 124–500 M models.
  • —Sigmoid router + grouped-bmm dispatch is the only supported backend at this scale; the legacy per-expert bmm backend would fall over on V100 throughput.
  • —router_noise_std / use_liger_ce / use_chunked_ce are forced false in config.json for inference; the training-time AMP / chunked-CE / Liger fused-CE paths are documented in the bundled model.py but are not exercised by load_test.py.