CoolFace
Modelpublic

jean-philippe-gire/Qwen3.8-Flash-Next-heretic-2-IQ4_XS-GGUF

sourceHugging Faceotherupdated 26d agoView on Hugging Face
0likes1.3kdownloads
Model Card

Qwen3.8-Flash-Next-heretic-2 — IQ4_XS GGUF

Minimal mirror containing the five plain IQ4XS GGUF shards, the six-file MTP/NextN view required for speculative decoding, and the matching Q80 multimodal projector. It is intended for Runpod Serverless host-side model caching, avoiding the unrelated quantizations in the full source repository.

Source: spiritfather/Qwen3.8-Flash-Next-heretic-2-i1-GGUF, pinned at revision 83e364306943b1466ecf21d97f8353e8a24a5dee.

Load the first shard with llama.cpp; the remaining shards are discovered automatically:

text
IQ4_XS/Qwen3.8-Flash-Next-heretic-2-IQ4_XS-00001-of-00005.gguf

For MTP, use a llama.cpp build containing PR #27836 and load:

text
IQ4_XS/MTP/Qwen3.8-Flash-Next-heretic-2-IQ4_XS-MTP-00001-of-00006.gguf

Recommended IQ4-class speculative settings:

text
--spec-type draft-mtp --spec-draft-n-max 7 --spec-draft-p-min 0.75

MTP shards 2–5 reference the same underlying Hugging Face LFS blobs as the plain trunk shards, so the cache does not need a second physical copy of those weights.

For image inputs, load the projector from the matching heretic-2 conversion:

text
mmproj/mmproj-Qwen3.8-Flash-Next-heretic-2-Q8_0.gguf

With llama.cpp, use --mmproj <path> --mmproj-offload. Qwen-VL grounding tasks benefit from --image-min-tokens 1024; this was validated on an RTX PRO 6000 Blackwell with the MTP shard set above.

The upstream model and Qwen Community License terms continue to apply.