dxx117/Qwen3.5-397B-A17B-GGUF-UD-IQ2_XXS-MTP-Q8_0
071
Qwen3.5-397B-A17B-GGUF Unsloth UD IQ2_XXS + Q8 MTP
This is a GGUF release of Qwen3.5-397B-A17B using the Unsloth UD IQ2_XXS quant as the trunk model, with a Q8_0 MTP draft block grafted on for speculative decoding experiments.
What this is
- Base trunk: Unsloth Qwen3.5-397B-A17B UD IQ2_XXS
- Added draft block: Q8_0 MTP
- Format: GGUF
- Output: split GGUF parts for llama.cpp-compatible workflows
Important provenance note
The MTP layers in this release were grafted from a vanilla Qwen3.5-397B source, because some downstream 397B variants retain MTP metadata/config but do not retain the actual MTP weights.
That means this model is best treated as an experimental MTP graft
Intended use
This release is mainly for:
- speculative decoding / draft model experiments
- comparing draft acceptance rates across grafted 397B variants
Caveats
- MTP acceptance rate may vary significantly depending on how well the grafted draft block matches the trunk model.
- This release does not imply guaranteed quality or speedup versus non-MTP variants.
- If you use multimodal workflows, note that this upload is for the text model GGUF, any mmproj file is separate as it is unsupported in llama.cpp.
- You might need to adjust fit-target parameters to load the draft model into VRAM
Performance
- I got about a 10-30% speedup running llama.cpp PR 22673 with partial offload
Files
This repo contains the split GGUF parts for:
Unsloth IQ2_XXS trunk + Q8_0 MTP graft
Recommended runtime notes
For fair comparisons, test with:
- the same llama.cpp build
- the same context length
- the same GPU offload settings
- the same cache / draft settings
Credits
- Qwen for the original model family
- Unsloth for the UD IQ2_XXS quant trunk
- buzz for the convert.py grafting script: https://gist.github.com/buzz/1c439684d5e3f36492ae9f64ef7e3f67
- MTP grafting and GGUF packaging by dxx117
