TeichAI/Qwen3.8-27B-Fable-Distill-GGUF
Qwen3.8-27B-Fable-Distill — GGUF

As always, big thank you to @nightmedia for the benchmarks
GGUF conversions of TeichAI/Qwen3.8-27B-Fable-Distill, a BF16 finetune of Qwen3.8-27B (base: Qwen/Qwen3.8-27B) trained with Unsloth + TRL.
The model was trained on a public set of chat and agent traces from Fable 5 as well as a much larger corpus of private personal Fable 5 data.
Converted with llama.cpp b6b4344e.
MTP head kept at BF16
This model ships a multi-token-prediction (nextn) head, and every quant here keeps that head unquantized at BF16 while the other 64 layers are quantized normally:
qwen35.block_count = 65 # 64 transformer layers + 1 MTP layer
qwen35.nextn_predict_layers = 1
blk.64.* = bf16 # 424.7M params, left aloneFiles
Usage
Text:
llama-cli -m Qwen3.8-27B-Fable-Distill-Q4_K_M.gguf -c 8192 -p "Hello"Vision — pass the projector alongside the model:
llama-mtmd-cli -m Qwen3.8-27B-Fable-Distill-Q4_K_M.gguf \
--mmproj mmproj-F16.gguf \
--image photo.jpg -p "Describe this image."Server:
llama-server -m Qwen3.8-27B-Fable-Distill-Q4_K_M.gguf --mmproj mmproj-F16.ggufServer + MTP:
llama-server -m Qwen3.8-27B-Fable-Distill-Q4_K_M.gguf --mmproj mmproj-F16.gguf --spec-type draft-mtp --spec-draft-n-max 3Notes
- The model is multimodal (image-text-to-text). Without an
mmproj-*.ggufyou get a text-only model. Three precisions are provided;F16is the usual choice,BF16matches the source weights' dtype, andF32is there if you want the projector left entirely unquantized. - Qwen3.5-family chat template with thinking support: it accepts
enable_thinkingand areasoning_effortoflow,mediumorxhigh(the template's own default isxhigh, which thinks at length every turn). - Base model sampling recommendations:
temperature 1.0,top_p 0.95,top_k 20.
The data for this model was easily formatted, validated, and masked using Teich <img src="https://cdn-avatars.huggingface.co/v1/production/uploads/6837935ac3b7ffe0d2559ce9/-AxyvV4wfUY8uo87kNKkK.png" width="20" height="20" style="display: inline-block; vertical-align: middle; margin: 0 3px;">
This qwen3_5 model was trained 2x faster with Unsloth and Huggingface's TRL library.
