CoolFace
Modelpublic

DAXZEIT/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled-UD-Q5_K_XL-gguf

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
1likes180downloads
Model Card

Qwen3.6-27B-Claude-Opus-Reasoning-Distilled — UD Q5KXL GGUF

Update 20/05/2026 — MTP draft head available: A compatible self-speculative decoding companion is now published: Qwen3.6-27B-Claude-Opus-Reasoning-MTP-Q4_K_M-gguf — load it via --model-draft for +40% decode throughput at no quality cost (llama.cpp ≥ b9245, --spec-type draft-mtp).

Quantized GGUF of rico03/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled in Unsloth Dynamic 2.0 Q5_K_XL format.

This is the only publicly available UD Q5KXL quantization of this model.

⚠️ Chat template fix required — the original rico03 repo ships the Qwen3.5 template, breaking Qwen3.6 tool calls. Add --chat-template-file qwen3.6_chat_template.jinja to your llama-server command (template source). Tool call format: <function=tool_name>{"param": "value"}</function>
For maximum context (212K on 24GB VRAM), see UD Q4_K_XL. For quality reference, see UD Q6_K_XL.

Part of the DAXZEIT UD SSM-aware series — UD recipes reverse-engineered from Unsloth base models and applied to distills using calibrated imatrix.

Updated 2026-05-12 — If you downloaded this model before 12 May 2026, please re-download. The initial release used an incorrect quantization recipe. This version applies the correct Unsloth Dynamic 2.0 UD recipe with calibrated imatrix.

Files

FileSizeBPWDescription
Qwen3.6-27B-Claude-Opus-Reasoning-Distilled-UD-Q5_K_XL.gguf19 GB6.04UD Q5KXL with imatrix

What is UD Q5KXL?

Unsloth Dynamic 2.0 is a mixed-precision quantization strategy that selectively upgrades the most accuracy-sensitive tensors above the base quantization level, guided by a per-tensor importance matrix (imatrix). The recipe is reverse-engineered from the official Unsloth base model GGUF and applied to the distill using a calibrated imatrix.

For Q5KXL the recipe uses 195 individual `--tensor-type` overrides on top of a Q5KM base:

Tensor distribution (851 total)

TypeCountTensors
Q8_048blk.N.ssm_out.weight — all SSM output projections
Q6_K176output.weight, attn_qkv, ffn_down, attn_v, late-network ffn_gate/up
Q5_K262attn_gate + remaining attention/FFN weights
Q4_K12Mid-network ffn_gate/up (blk.12–16, imatrix-guided)
F32353Norms, biases, SSM scalars

The 48 ssm_out.weight tensors (one per block) are kept at Q8_0 — these are the output projections of the SSM recurrent mechanism, critical for long-context coherence in the Qwen3.6 hybrid architecture.


Perplexity

Measured on wikitext-2, full test set (~1M tokens, 1952 chunks, ctx=512), RTX 3090.

ModelPPLBPWSize
Q6_K plain (reference)7.4694 ±0.0316.5721 GB
UD Q4KXL7.4712 ±0.0315.4117 GB
UD Q5_K_XL (this)7.4891 ±0.0316.0419 GB
UD Q6KXL7.4753 ±0.0317.6424 GB

All four models fall within 0.02 PPL of each other — statistically equivalent (±σ overlap on all). The UD recipe concentrates precision on the tensors that matter most, achieving Q6_K-level quality at a lower bit-per-weight budget.


Vision (mmproj)

This model supports vision via a multimodal projector compatible with the Qwen3.6-27B hybrid SSM+Transformer architecture.

mmproj: DAXZEIT/Qwen3.6-27B-mmproj-hybrid-Q8_0-F16-gguf — 601 MB, Q8_0/F16 mixed precision
bash
llama-server \
    -m <this-model>.gguf \
    --mmproj Qwen3.6-27B-mmproj-hybrid-Q8_0-F16.gguf \
    ...

Usage

bash
llama-server \
    -m Qwen3.6-27B-Claude-Opus-Reasoning-Distilled-UD-Q5_K_XL.gguf \
    --mmproj Qwen3.6-27B-mmproj-Q8_0.gguf \
    --ctx-size 131072 \
    --n-gpu-layers 65 \
    --cache-type-k q4_0 \
    --cache-type-v q4_0 \
    --flash-attn auto \
    --port 5000

Fits in 24GB VRAM at 65 GPU layers, Q4_0 KV cache, 131K context with mmproj.


Quantization Recipe

imatrix source: rico03/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled Q6_K — wikitext-2, 200 chunks, ctx 512.

The UD recipe was reverse-engineered from unsloth/Qwen3.6-27B-GGUF UD-Q5_K_XL using GGUFReader tensor extraction. The architecture is identical between base model and distill, so the tensor override map transfers directly.

python
# Extract recipe from reference GGUF
from gguf import GGUFReader
r = GGUFReader('unsloth/Qwen3.6-27B-UD-Q5_K_XL.gguf')
for t in r.tensors:
    if t.tensor_type.name not in ('Q5_K', 'F32'):
        print(f'--tensor-type {t.name}={t.tensor_type.name}')
# → 195 overrides
bash
llama-quantize \
    --imatrix Qwen3.6-27B-Claude-Opus-Reasoning-Distilled.imatrix \
    [195 --tensor-type overrides] \
    rico03-distill-f16.gguf \
    output-UD-Q5_K_XL.gguf \
    Q5_K_M

Architecture

Qwen3.6-27B hybrid SSM+Transformer:

  • —64 blocks — each block contains both SSM and attention components
  • —Context: 262K tokens native
  • —Vocabulary: 248,320 tokens

Credits


License

Apache 2.0 — inherited from the base model.