DAXZEIT/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled-UD-Q5_K_XL-gguf
Qwen3.6-27B-Claude-Opus-Reasoning-Distilled — UD Q5KXL GGUF
Update 20/05/2026 — MTP draft head available: A compatible self-speculative decoding companion is now published: Qwen3.6-27B-Claude-Opus-Reasoning-MTP-Q4_K_M-gguf — load it via--model-draftfor +40% decode throughput at no quality cost (llama.cpp ≥ b9245,--spec-type draft-mtp).
Quantized GGUF of rico03/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled in Unsloth Dynamic 2.0 Q5_K_XL format.
This is the only publicly available UD Q5KXL quantization of this model.
⚠️ Chat template fix required — the original rico03 repo ships the Qwen3.5 template, breaking Qwen3.6 tool calls. Add--chat-template-file qwen3.6_chat_template.jinjato your llama-server command (template source). Tool call format:<function=tool_name>{"param": "value"}</function>
For maximum context (212K on 24GB VRAM), see UD Q4_K_XL. For quality reference, see UD Q6_K_XL.
Part of the DAXZEIT UD SSM-aware series — UD recipes reverse-engineered from Unsloth base models and applied to distills using calibrated imatrix.
Updated 2026-05-12 — If you downloaded this model before 12 May 2026, please re-download. The initial release used an incorrect quantization recipe. This version applies the correct Unsloth Dynamic 2.0 UD recipe with calibrated imatrix.
Files
What is UD Q5KXL?
Unsloth Dynamic 2.0 is a mixed-precision quantization strategy that selectively upgrades the most accuracy-sensitive tensors above the base quantization level, guided by a per-tensor importance matrix (imatrix). The recipe is reverse-engineered from the official Unsloth base model GGUF and applied to the distill using a calibrated imatrix.
For Q5KXL the recipe uses 195 individual `--tensor-type` overrides on top of a Q5KM base:
Tensor distribution (851 total)
The 48 ssm_out.weight tensors (one per block) are kept at Q8_0 — these are the output projections of the SSM recurrent mechanism, critical for long-context coherence in the Qwen3.6 hybrid architecture.
Perplexity
Measured on wikitext-2, full test set (~1M tokens, 1952 chunks, ctx=512), RTX 3090.
All four models fall within 0.02 PPL of each other — statistically equivalent (±σ overlap on all). The UD recipe concentrates precision on the tensors that matter most, achieving Q6_K-level quality at a lower bit-per-weight budget.
Vision (mmproj)
This model supports vision via a multimodal projector compatible with the Qwen3.6-27B hybrid SSM+Transformer architecture.
mmproj: DAXZEIT/Qwen3.6-27B-mmproj-hybrid-Q8_0-F16-gguf — 601 MB, Q8_0/F16 mixed precision
llama-server \
-m <this-model>.gguf \
--mmproj Qwen3.6-27B-mmproj-hybrid-Q8_0-F16.gguf \
...Usage
llama-server \
-m Qwen3.6-27B-Claude-Opus-Reasoning-Distilled-UD-Q5_K_XL.gguf \
--mmproj Qwen3.6-27B-mmproj-Q8_0.gguf \
--ctx-size 131072 \
--n-gpu-layers 65 \
--cache-type-k q4_0 \
--cache-type-v q4_0 \
--flash-attn auto \
--port 5000Fits in 24GB VRAM at 65 GPU layers, Q4_0 KV cache, 131K context with mmproj.
Quantization Recipe
imatrix source: rico03/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled Q6_K — wikitext-2, 200 chunks, ctx 512.
The UD recipe was reverse-engineered from unsloth/Qwen3.6-27B-GGUF UD-Q5_K_XL using GGUFReader tensor extraction. The architecture is identical between base model and distill, so the tensor override map transfers directly.
# Extract recipe from reference GGUF
from gguf import GGUFReader
r = GGUFReader('unsloth/Qwen3.6-27B-UD-Q5_K_XL.gguf')
for t in r.tensors:
if t.tensor_type.name not in ('Q5_K', 'F32'):
print(f'--tensor-type {t.name}={t.tensor_type.name}')
# → 195 overridesllama-quantize \
--imatrix Qwen3.6-27B-Claude-Opus-Reasoning-Distilled.imatrix \
[195 --tensor-type overrides] \
rico03-distill-f16.gguf \
output-UD-Q5_K_XL.gguf \
Q5_K_MArchitecture
Qwen3.6-27B hybrid SSM+Transformer:
- 64 blocks — each block contains both SSM and attention components
- Context: 262K tokens native
- Vocabulary: 248,320 tokens
Credits
- Base model: rico03/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled
- imatrix: generated from rico03 Q6_K on wikitext-2
- UD recipe: reverse-engineered from unsloth/Qwen3.6-27B-GGUF
- Quantization methodology: Unsloth Dynamic 2.0
- Quantized by: DAXZEIT
License
Apache 2.0 — inherited from the base model.
