DAXZEIT/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled-UD-Q4_K_XL-gguf
Qwen3.6-27B-Claude-Opus-Reasoning-Distilled — UD Q4KXL GGUF
Update 20/05/2026 — MTP draft head available: A compatible self-speculative decoding companion is now published: Qwen3.6-27B-Claude-Opus-Reasoning-MTP-Q4_K_M-gguf — load it via--model-draftfor +40% decode throughput at no quality cost (llama.cpp ≥ b9245,--spec-type draft-mtp).
Quantized GGUF of rico03/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled in Unsloth Dynamic 2.0 Q4_K_XL format.
Use case: the only quantization of this model that fits the full native 262K context on a single 24GB GPU — confirmed on RTX 3090.
llama-server + model + KV cache (262K, Q4_0) : 22,800 MiB / 24,576 MiBJust fits. For maximum reasoning quality at shorter context, see UD Q5_K_XL.
⚠️ Chat template fix required — the original rico03 repo ships the Qwen3.5 template, breaking Qwen3.6 tool calls. Add--chat-template-file qwen3.6_chat_template.jinjato your llama-server command (template source). Tool call format:<function=tool_name>{"param": "value"}</function>
Part of the DAXZEIT UD SSM-aware series — UD recipes reverse-engineered from Unsloth base models and applied to distills using calibrated imatrix.
Updated 2026-05-12 — If you downloaded this model before 12 May 2026, please re-download. The initial release used an incorrect quantization recipe. This version applies the correct Unsloth Dynamic 2.0 UD recipe with calibrated imatrix.
Files
Not your average Q4
Despite the name, this is not a plain Q4KM. The effective BPW is 5.41 — higher than a standard Q5KM (5.00 BPW) — because the UD recipe selectively promotes critical tensors well above Q4.
Tensor distribution (851 total)
Compared to a Q4KM vanilla (all ~498 weight tensors at Q4K): **more than half the tensors are above Q4**, with the 48 SSM outputs at full Q80.
The 12 mid-network ffn_gate/up tensors (blk.12–16) that appear as IQ4XS in the Unsloth base model are **promoted to Q5K** in this distill — the imatrix indicates these paths are more active in the Opus reasoning fine-tune.
Perplexity
Measured on wikitext-2, full test set (~1M tokens, 1952 chunks, ctx=512), RTX 3090.
All four models fall within 0.02 PPL of each other — statistically equivalent (±σ overlap on all). The UD Q4KXL at 5.41 BPW matches the plain Q6_K at 6.57 BPW, saving 4 GB with no measurable quality loss on this benchmark.
Vision (mmproj)
This model supports vision via a multimodal projector compatible with the Qwen3.6-27B hybrid SSM+Transformer architecture.
mmproj: DAXZEIT/Qwen3.6-27B-mmproj-hybrid-Q8_0-F16-gguf — 601 MB, Q8_0/F16 mixed precision
llama-server \
-m <this-model>.gguf \
--mmproj Qwen3.6-27B-mmproj-hybrid-Q8_0-F16.gguf \
...Usage
llama-server \
-m Qwen3.6-27B-Claude-Opus-Reasoning-Distilled-UD-Q4_K_XL.gguf \
--ctx-size 262144 \
--n-gpu-layers 65 \
--cache-type-k q4_0 \
--cache-type-v q4_0 \
--flash-attn auto \
--port 5000Note:--flash-attn autowith--cache-type-k q4_0requires llama.cpp compiled withGGML_CUDA_FA_ALL_QUANTS=ON, otherwise flash attention silently falls back to standard attention on quantized KV types.
Confirmed on RTX 3090 — 22,800 MiB / 24,576 MiB at 262K native context. Tight fit, but stable.
Quantization Recipe
imatrix source: rico03/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled Q6_K — wikitext-2, 200 chunks, ctx 512.
The UD recipe was reverse-engineered from unsloth/Qwen3.6-27B-GGUF UD-Q4_K_XL using GGUFReader tensor extraction. The architecture is identical between base model and distill, so the tensor override map transfers directly.
# Extract recipe from reference GGUF
from gguf import GGUFReader
r = GGUFReader('unsloth/Qwen3.6-27B-UD-Q4_K_XL.gguf')
for t in r.tensors:
if t.tensor_type.name not in ('Q4_K', 'F32'):
print(f'--tensor-type {t.name}={t.tensor_type.name}')
# → 195 overridesllama-quantize \
--imatrix Qwen3.6-27B-Claude-Opus-Reasoning-Distilled.imatrix \
[195 --tensor-type overrides] \
rico03-distill-f16.gguf \
output-UD-Q4_K_XL.gguf \
Q4_K_MArchitecture
Qwen3.6-27B hybrid SSM+Transformer:
- 64 blocks — each block contains both SSM and attention components
- Context: 262K tokens native
- Vocabulary: 248,320 tokens
Credits
- Base model: rico03/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled
- imatrix: generated from rico03 Q6_K on wikitext-2
- UD recipe: reverse-engineered from unsloth/Qwen3.6-27B-GGUF
- Quantization methodology: Unsloth Dynamic 2.0
- Quantized by: DAXZEIT
License
Apache 2.0 — inherited from the base model.
