gsrunion/Qwen3.6-35B-A3B-ROCmFP4-STRIX_LEAN-DFLASH-GGUF
Qwen3.6-35B-A3B — ROCmFP4 STRIX_LEAN, DFlash baked in
A single-file, self-accelerating GGUF: the model and its DFlash speculative-decoding draft are merged into one .gguf. No --model-draft, no --spec-type flag — point -m at this file and speculative decoding just happens.
To our knowledge, the first "draft-included" GGUF publication anywhere.
llama-server -m Qwen3.6-35B-A3B-STRIX_LEAN-DFLASH.gguf -ngl 999 -fa on --jinja -c 65536Requirements
This needs both a ROCmFP4-aware build and the DFlash-graft support for embedded drafts — neither exists upstream yet. Use:
- gsrunion/rocmfp4-llama branch
dflash-graft(built and validated on AMD Strix Halo / gfx1151), or - your own build once the fixes below land in charlie12345/rocmfp4-llama
On first load the server extracts the draft's tensors to a small cached sidecar file next to the model (one-time, ~1 second).
Measured performance
AMD Ryzen AI Max+ 395 (Strix Halo, 128 GB unified LPDDR5X), server-timing, self-accelerating load (zero extra flags):
Within noise of the two-file config — the merge adds no overhead.
How it was made
The draft's tensors are merged into the target GGUF prefixed dflash.* (target keeps its own tensor names untouched — no collision, no size overhead: DFlash drafts already borrow the target's token embeddings and output head at runtime, so nothing is duplicated). A dflash.embedded marker key flags the file for auto-detection.
Two fixes were needed in the serving fork to make this work (both filed against the base fork, worth watching if you hit similar issues building your own):
- The tensor-count sanity check in the model loader didn't allow "extra" tensors belonging to a sibling model in the same file — even though the check already had unused plumbing for exactly this case.
- The draft's
mask_token_id(namespaced undertokenizer.*by convention, though it's actually draft-specific) has to be copied into the merged file explicitly, or drafting silently no-ops with zero speedup and no error.
Base weights: gsrunion/Qwen3.6-35B-A3B-ROCmFP4-STRIX_LEAN-GGUF. Draft: z-lab/Qwen3.6-35B-A3B-DFlash.
Credits
- Base model: Qwen — Qwen3.6-35B-A3B (Apache-2.0)
- DFlash draft: z-lab
- ROCmFP4 quant formats: Hal0ai; fork base: charlie12345
