ranxianglei/Qwen3.8-Flash-Next-W4A16-Modular
Qwen3.8-Flash-Next W4A16 Modular (296E)
Expert-pruned, expert-modular repack of the Intel AutoRound W4A16 build of Qwen3.8-Flash-Next: 512 → 296 experts per layer, selected by profiling real agent traffic. Serves on a single 96 GB GPU with a 262K context window.
How this was made
- Base: official Qwen/Qwen3.8-Flash-Next → Intel AutoRound W4A16 (calibrated int4).
- Profile: ~80 real coding-agent sessions were replayed against a server running SGLang's
--expert-distribution-recorder-mode stat, producing per-(layer, expert) route counts. A second profile was taken on anomalous contexts (the failure sessions we wanted the model to keep handling well). - Keep set: per layer, top-294 experts from daily traffic ∪ top-2 from the anomaly profile → 296/layer. Covers 95%+ of routine routing and preserves self-healing on degraded contexts (measured 5/5 recovery vs 0/3 for the daily-only set).
- Modular repack: all expert tensors of layer N were re-packed into one
experts-L{NN}.safetensors; everything else (dense, GDN linear attention, PLE, embeddings, lm_head) into 14 large backbone shards. No tensor values were modified — this is a pure re-chunking of the same weights, verified key-for-key identical (129,403 tensors) with serving parity (104 tok/s).
Why it works (principle)
MoE layers route each token to only top-k of num_experts experts (here 10 of 512). Routine traffic concentrates on a small subset per layer, so removing never-routed experts is lossless for that workload; the anomaly-profile union buys back robustness for edge contexts. Expert weights live in GPU memory only when kept, so pruning 512→296 frees ~13 GB VRAM → +31% KV pool.
Changes vs the previous release (Pruned-294E)
Same base weights, same quantization; day-to-day quality and speed are identical.
Modular layout
Swapping, adding or re-pruning experts for layer N only rewrites experts-LNN.safetensors plus config.json (num_experts) and the index weight_map — the backbone never changes.
Want a keep set tuned to your traffic? Profile it in one command and serve pruned without re-exporting: sglang-expert-profile (CLI + community keep-sets; the serving-side keep-mask lives in the sglang fork, ours/main).
Requirements
- GPU VRAM >= 64 GB (weights ~45 GB; 96 GB recommended for full 262K context)
- Host RAM >= 64 GB (PLE embedding offload:
--ple-offload-embedding) - CUDA 13 stack
Serving (SGLang fork with PLE offload + marlin GC fix)
python -m sglang.launch_server \
--model-path ./Qwen3.8-Flash-Next-W4A16-Modular \
--chat-template ./qwen3_coder_template.jinja \
--ple-offload-embedding \
--moe-a2a-backend none \
--linear-attn-prefill-backend triton --linear-attn-decode-backend triton \
--mamba-ssm-dtype bfloat16 \
--context-length 262144 --mem-fraction-static 0.93 \
--reasoning-parser qwen3 --tool-call-parser qwen3_coderSampling defaults ship in generation_config.json (temp 0.7 / topp 0.95 / topk 20, no penalties — Flash-Next is penalty-sensitive, see our notes in the sglang fork). No need to pass them per request.
Note: without the marlin GC patch, loading OOMs at ~91.5 GB on some stacks (gptqmarlinmoe_repack int4→int32 expansion). Patch + details: https://github.com/ranxianglei/sglang (ours branch).
Performance (single RTX Pro 6000 96GB)
- single stream ~104 tok/s decode @ 262K context
- aggregate (w48) ~2100 tok/s
- KV pool: ~856K tokens bf16
Pairs well with billion-context (ACP)
This model's 262K window + huge KV pool makes it an excellent host for our context-compression plugins — long agent sessions stay coherent while effective context grows far beyond the window:
- billion-context — protocol-level ACP context compression for AI coding agents (drop-in for OpenCode & friends)
- billion-context-pi — pi/agent integration of the same compression engine
Together: Flash-Next serves the window, billion-context compresses into it — day-long coding agents on one consumer GPU.
Provenance
Base: Qwen/Qwen3.8-Flash-Next → Intel AutoRound W4A16 (calibrated) → expert pruning 512 → 296 by routing-profile keep-set → modular repack (this repo). 294-expert variant (linear shards): ranxianglei/Qwen3.8-Flash-Next-W4A16-Pruned-294E
