HYHPING2023/checkpoint-draft-dflash2
checkpoint-draft-dflash2 (Qwen3.5-35B-A3B VCLR3 drafter)
DFlash2 speculative-decoding draft model for the Qwen3.5-35B-A3B (VCLR3 game-video SFT) target. Not a standalone language model: it runs inside a speculative decoding server and drafts tokens for the target to verify (target weights NOT included; point --model-path at your own merged Qwen3.5-35B-A3B VCLR3 checkpoint).
Architecture
DFlash2 (Inco AI, blog; z-lab-compatible weight layout) with block_size 8:
- 5-layer Qwen3-style dual-stream draft over the target's captured layer
{1, 10, 19, 28, 37}hidden states, full 248k vocab (uses the target's embedtokens / lmhead) - Two-tap grouped dynamic causal convolutions (kernel 2, group 16) around every attention/MLP sublayer — fixes block-end (suffix) decay
- Top-16 candidate path selector (rank 256):
S_t(a,b) = U_t(b) + <A(a) ⊙ H(h_t), B(b)>
Serving (SGLang main, DFlash2-aware DFLASH worker)
python -m sglang.launch_server \
--model-path /path/to/Qwen3.5-35B-A3B-vclr3-target \
--speculative-algorithm DFLASH \
--speculative-draft-model-path HYHPING2023/checkpoint-draft-dflash2 \
--speculative-num-draft-tokens 8 \
--trust-remote-code --tp 1 --enable-metricsNotes: request chat_template_kwargs: {"enable_thinking": false} (the draft was trained on non-thinking answers); SGLang builds before DFlash2 support silently ignore the selector.
Evaluation (temperature 0, greedy; block 8)
Per-position conditional acceptance rises 0.63 → 0.72 toward the block end (DFlash1 stays flat ~0.60) — the convolution's suffix-decay fix.
Training
4 epochs on the mixed video+zh+en corpus (70.6k samples), warm-started from the DFlash1 b8 mixed2 checkpoint, FSDP + frozen online 35B target. Teacher-forced selector CE. Checkpoint: epoch_3_step_35300.
