jean-philippe-gire/Qwen3.8-Flash-Next-heretic-2-IQ4_XS-GGUF
Qwen3.8-Flash-Next-heretic-2 — IQ4_XS GGUF
Minimal mirror containing the five plain IQ4XS GGUF shards, the six-file MTP/NextN view required for speculative decoding, and the matching Q80 multimodal projector. It is intended for Runpod Serverless host-side model caching, avoiding the unrelated quantizations in the full source repository.
Source: spiritfather/Qwen3.8-Flash-Next-heretic-2-i1-GGUF, pinned at revision 83e364306943b1466ecf21d97f8353e8a24a5dee.
Load the first shard with llama.cpp; the remaining shards are discovered automatically:
IQ4_XS/Qwen3.8-Flash-Next-heretic-2-IQ4_XS-00001-of-00005.ggufFor MTP, use a llama.cpp build containing PR #27836 and load:
IQ4_XS/MTP/Qwen3.8-Flash-Next-heretic-2-IQ4_XS-MTP-00001-of-00006.ggufRecommended IQ4-class speculative settings:
--spec-type draft-mtp --spec-draft-n-max 7 --spec-draft-p-min 0.75MTP shards 2–5 reference the same underlying Hugging Face LFS blobs as the plain trunk shards, so the cache does not need a second physical copy of those weights.
For image inputs, load the projector from the matching heretic-2 conversion:
mmproj/mmproj-Qwen3.8-Flash-Next-heretic-2-Q8_0.ggufWith llama.cpp, use --mmproj <path> --mmproj-offload. Qwen-VL grounding tasks benefit from --image-min-tokens 1024; this was validated on an RTX PRO 6000 Blackwell with the MTP shard set above.
The upstream model and Qwen Community License terms continue to apply.
