CoolFace
Modelpublic

gsrunion/Qwen3.6-35B-A3B-ROCmFP4-STRIX_LEAN-DFLASH-GGUF

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes342downloads
Model Card

Qwen3.6-35B-A3B — ROCmFP4 STRIX_LEAN, DFlash baked in

A single-file, self-accelerating GGUF: the model and its DFlash speculative-decoding draft are merged into one .gguf. No --model-draft, no --spec-type flag — point -m at this file and speculative decoding just happens.

To our knowledge, the first "draft-included" GGUF publication anywhere.

bash
llama-server -m Qwen3.6-35B-A3B-STRIX_LEAN-DFLASH.gguf -ngl 999 -fa on --jinja -c 65536

Requirements

This needs both a ROCmFP4-aware build and the DFlash-graft support for embedded drafts — neither exists upstream yet. Use:

On first load the server extracts the draft's tensors to a small cached sidecar file next to the model (one-time, ~1 second).

Measured performance

AMD Ryzen AI Max+ 395 (Strix Halo, 128 GB unified LPDDR5X), server-timing, self-accelerating load (zero extra flags):

tok/sacceptance
Baked single-file91.898.5% (405/411)
Two-file (--model-draft + flags)96.097–98%
Plain LEAN, no draft63.1—

Within noise of the two-file config — the merge adds no overhead.

How it was made

The draft's tensors are merged into the target GGUF prefixed dflash.* (target keeps its own tensor names untouched — no collision, no size overhead: DFlash drafts already borrow the target's token embeddings and output head at runtime, so nothing is duplicated). A dflash.embedded marker key flags the file for auto-detection.

Two fixes were needed in the serving fork to make this work (both filed against the base fork, worth watching if you hit similar issues building your own):

  1. 1.The tensor-count sanity check in the model loader didn't allow "extra" tensors belonging to a sibling model in the same file — even though the check already had unused plumbing for exactly this case.
  2. 2.The draft's mask_token_id (namespaced under tokenizer.* by convention, though it's actually draft-specific) has to be copied into the merged file explicitly, or drafting silently no-ops with zero speedup and no error.

Base weights: gsrunion/Qwen3.6-35B-A3B-ROCmFP4-STRIX_LEAN-GGUF. Draft: z-lab/Qwen3.6-35B-A3B-DFlash.

Credits