CoolFace
Modelpublic

DAXZEIT/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled-UD-Q4_K_XL-gguf

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
6likes307downloads
Model Card

Qwen3.6-27B-Claude-Opus-Reasoning-Distilled — UD Q4KXL GGUF

Update 20/05/2026 — MTP draft head available: A compatible self-speculative decoding companion is now published: Qwen3.6-27B-Claude-Opus-Reasoning-MTP-Q4_K_M-gguf — load it via --model-draft for +40% decode throughput at no quality cost (llama.cpp ≥ b9245, --spec-type draft-mtp).

Quantized GGUF of rico03/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled in Unsloth Dynamic 2.0 Q4_K_XL format.

Use case: the only quantization of this model that fits the full native 262K context on a single 24GB GPU — confirmed on RTX 3090.

llama-server + model + KV cache (262K, Q4_0) : 22,800 MiB / 24,576 MiB

Just fits. For maximum reasoning quality at shorter context, see UD Q5_K_XL.

⚠️ Chat template fix required — the original rico03 repo ships the Qwen3.5 template, breaking Qwen3.6 tool calls. Add --chat-template-file qwen3.6_chat_template.jinja to your llama-server command (template source). Tool call format: <function=tool_name>{"param": "value"}</function>

Part of the DAXZEIT UD SSM-aware series — UD recipes reverse-engineered from Unsloth base models and applied to distills using calibrated imatrix.

Updated 2026-05-12 — If you downloaded this model before 12 May 2026, please re-download. The initial release used an incorrect quantization recipe. This version applies the correct Unsloth Dynamic 2.0 UD recipe with calibrated imatrix.

Files

FileSizeBPWDescription
Qwen3.6-27B-Claude-Opus-Reasoning-Distilled-UD-Q4_K_XL.gguf17 GB5.41UD Q4KXL with imatrix

Not your average Q4

Despite the name, this is not a plain Q4KM. The effective BPW is 5.41 — higher than a standard Q5KM (5.00 BPW) — because the UD recipe selectively promotes critical tensors well above Q4.

Tensor distribution (851 total)

TypeCountTensors
Q8_048blk.N.ssm_out.weight — all SSM output projections
Q6_K110output.weight, attn_qkv, ffn_down, attn_v (critical paths)
Q5_K70attn_gate + imatrix-promoted weights (incl. mid-network ffn_gate/up)
Q4_K258Remaining attention/FFN weights
F32353Norms, biases, SSM scalars

Compared to a Q4KM vanilla (all ~498 weight tensors at Q4K): **more than half the tensors are above Q4**, with the 48 SSM outputs at full Q80.

The 12 mid-network ffn_gate/up tensors (blk.12–16) that appear as IQ4XS in the Unsloth base model are **promoted to Q5K** in this distill — the imatrix indicates these paths are more active in the Opus reasoning fine-tune.


Perplexity

Measured on wikitext-2, full test set (~1M tokens, 1952 chunks, ctx=512), RTX 3090.

ModelPPLBPWSize
Q6_K plain (reference)7.4694 ±0.0316.5721 GB
UD Q4_K_XL (this)7.4712 ±0.0315.4117 GB
UD Q5KXL7.4891 ±0.0316.0419 GB
UD Q6KXL7.4753 ±0.0317.6424 GB

All four models fall within 0.02 PPL of each other — statistically equivalent (±σ overlap on all). The UD Q4KXL at 5.41 BPW matches the plain Q6_K at 6.57 BPW, saving 4 GB with no measurable quality loss on this benchmark.


Vision (mmproj)

This model supports vision via a multimodal projector compatible with the Qwen3.6-27B hybrid SSM+Transformer architecture.

mmproj: DAXZEIT/Qwen3.6-27B-mmproj-hybrid-Q8_0-F16-gguf — 601 MB, Q8_0/F16 mixed precision
bash
llama-server \
    -m <this-model>.gguf \
    --mmproj Qwen3.6-27B-mmproj-hybrid-Q8_0-F16.gguf \
    ...

Usage

bash
llama-server \
    -m Qwen3.6-27B-Claude-Opus-Reasoning-Distilled-UD-Q4_K_XL.gguf \
    --ctx-size 262144 \
    --n-gpu-layers 65 \
    --cache-type-k q4_0 \
    --cache-type-v q4_0 \
    --flash-attn auto \
    --port 5000
Note: --flash-attn auto with --cache-type-k q4_0 requires llama.cpp compiled with GGML_CUDA_FA_ALL_QUANTS=ON, otherwise flash attention silently falls back to standard attention on quantized KV types.

Confirmed on RTX 3090 — 22,800 MiB / 24,576 MiB at 262K native context. Tight fit, but stable.


Quantization Recipe

imatrix source: rico03/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled Q6_K — wikitext-2, 200 chunks, ctx 512.

The UD recipe was reverse-engineered from unsloth/Qwen3.6-27B-GGUF UD-Q4_K_XL using GGUFReader tensor extraction. The architecture is identical between base model and distill, so the tensor override map transfers directly.

python
# Extract recipe from reference GGUF
from gguf import GGUFReader
r = GGUFReader('unsloth/Qwen3.6-27B-UD-Q4_K_XL.gguf')
for t in r.tensors:
    if t.tensor_type.name not in ('Q4_K', 'F32'):
        print(f'--tensor-type {t.name}={t.tensor_type.name}')
# → 195 overrides
bash
llama-quantize \
    --imatrix Qwen3.6-27B-Claude-Opus-Reasoning-Distilled.imatrix \
    [195 --tensor-type overrides] \
    rico03-distill-f16.gguf \
    output-UD-Q4_K_XL.gguf \
    Q4_K_M

Architecture

Qwen3.6-27B hybrid SSM+Transformer:

  • —64 blocks — each block contains both SSM and attention components
  • —Context: 262K tokens native
  • —Vocabulary: 248,320 tokens

Credits


License

Apache 2.0 — inherited from the base model.