CoolFace
Modelpublic

lambsea/Qwen3.6-27B-AEON-Ultimate-Uncensored-UD-GGUF

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
1likes233downloads
Model Card

Qwen3.6-27B-AEON-Ultimate-Uncensored — GGUF (UD Quants)

Unsloth Dynamic-style (UD) GGUF quantizations of AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16

Every quant uses per-tensor overrides (sensitivity-driven) + importance matrix (multi-domain calibration). All SSM recurrence tensors are preserved at source precision. MTP speculative decoding and vision (mmproj) are preserved.


Quant Comparison

FileQuantSizetg t/sPPLKL meanKL maxKL p99.9
F16F1650.9 GB30.95.7215———
UD-Q8_0Q8_034.7 GB44.75.70550.00295.230.22
UD-Q6_KQ6_K30.6 GB49.15.70460.00457.640.33
UD-Q5KMQ5KM28.7 GB46.95.75890.01115.092.04
UD-IQ4_XSIQ4_XS25.9 GB56.55.76300.02364.531.97

Benchmarked on NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM), llama.cpp fork (a4501150/llama.cpp), pp=512, tg=128.


What Makes These Different

SSM Recurrence Preservation

Qwen3.6 is a hybrid GatedDeltaNet + attention model. 48 of 64 layers use a recurrent SSM where quantization error compounds across token positions. All SSM recurrence tensors are preserved at source precision (F16) — never quantized.

TensorCountPrecisionRationale
ssm_alpha, ssm_beta96F16State update projections — error accumulates in recurrence
ssm_out48F16Output projection feeds directly into residual stream
ssm_a, ssm_conv1d, ssm_dt, ssm_norm192F32Small state tensors (llama-quantize keeps 1D/small tensors at F32)
attn_qkv (SSM input projection)48F16Highest measured KL sensitivity
attn_gate (SSM gate projection)48F16Second-highest measured KL sensitivity

Per-Tensor Sensitivity Analysis

Each tensor group was probed by quantizing only that group to Q4_0 while keeping the rest at F16, then measuring KL divergence. The override generator assigns precision based on measured sensitivity:

PrecisionTensor GroupsOverride Count
F16SSM recurrence, norms, biases, MTP layer512
F16All attention tensors (attn_qkv, attn_gate, attn_v, attn_q, attn_k, attn_output), ffn_down edge173
Base quantFFN middle layers, FFN edge gate/up, embeddings~181

685 total overrides — every non-FFN tensor has an explicit precision assignment. No dependence on llama-quantize's internal promotion rules.

Multi-Domain Calibration + GPU Imatrix

Calibrated on a balanced mix across 4 domains from 13 HF datasets:

DomainToken BudgetSources
General1Multrachat, OpenHermes, COIG-CQIA, LongAlpaca, pg19, froggeric/imatrix
Code750KMagicoder-Evol-Instruct-110K
Reasoning750KOpenMathInstruct-2, OpenR1-Math-220k
Agentic500Kglaive-function-calling-v2, xlam-function-calling-60k, hermes-function-calling-v1

Special tokens from source datasets are stripped automatically. Samples are kept whole — never truncated mid-conversation.

The importance matrix is generated with a PyTorch GPU-native generator (src/generate_imatrix.py) at 65,536 context — uses forward hooks to accumulate squared activations on GPU with zero PCIe D2H copies. Supports multi-GPU via device_map="auto".

Per-domain imatrices are merged with equal weights (DI-MATRIX approach).

MTP + Vision Preserved

  • —MTP (Multi-Token Prediction): Draft head (blk.64) pinned at F16. Use --spec-type draft-mtp --spec-draft-n-max 3 for ~1.5-2x faster generation.
  • —Vision: mmproj file contains the full vision encoder. Use --mmproj flag with llama-server for image/video understanding.

Files

FileDescriptionSize
Qwen3.6-27B-AEON-UD-Q8_0.ggufHighest quality quantization34.7 GB
Qwen3.6-27B-AEON-UD-Q6_K.ggufRecommended — best quality/size30.6 GB
Qwen3.6-27B-AEON-UD-Q5_K_M.ggufBalanced28.7 GB
Qwen3.6-27B-AEON-UD-IQ4_XS.ggufSmallest, for constrained VRAM25.9 GB
Qwen3.6-27B-AEON-mmproj-F16.ggufVision encoder (use with --mmproj)885 MB
imatrix_merged.datImportance matrix for requantization13 MB

Usage

llama-server (recommended)

bash
# Q6_K with YaRN 512k context, 5 concurrent slots, MTP + vision
llama-server \
    -m Qwen3.6-27B-AEON-UD-Q6_K.gguf \
    --mmproj Qwen3.6-27B-AEON-mmproj-F16.gguf \
    -ngl 99 \
    --flash-attn \
    -c 524288 \
    --parallel 5 \
    --cache-type-k q8_0 \
    --cache-type-v q8_0 \
    -kvu \
    --cache-ram -1 \
    --rope-scaling yarn \
    --rope-scale 2.0 \
    --yarn-orig-ctx 262144 \
    --override-kv "qwen35.context_length=int:524288" \
    --spec-type draft-mtp \
    --spec-draft-n-max 3 \
    --jinja \
    --chat-template-kwargs '{"enable_thinking":true,"preserve_thinking":true}' \
    --host 0.0.0.0 --port 8080
Note: --spec-type draft-mtp requires llama.cpp b9375+. A custom fork adds DFlash speculative decoding and Blackwell-tuned flash attention.

llama-cli

bash
llama-cli \
    -m Qwen3.6-27B-AEON-UD-Q6_K.gguf \
    -ngl 99 \
    --flash-attn \
    -c 524288 \
    --rope-scaling yarn \
    --rope-scale 2.0 \
    --yarn-orig-ctx 262144 \
    --jinja \
    --chat-template-kwargs '{"enable_thinking":true,"preserve_thinking":true}'

Chat Template Notes

  • —enable_thinking activates reasoning mode (chain-of-thought in <think> blocks)
  • —preserve_thinking retains reasoning blocks in conversation history
  • —No spaces after colons in the JSON — Qwen3.6's template parser is whitespace-sensitive

Architecture

Qwen3.6-27B is a hybrid SSM-attention model:

  • —64 transformer layers + 1 MTP layer (blk.0-64)
  • —48 SSM layers (GatedDeltaNet, no KV cache) + 16 full attention layers (every 4th layer)
  • —27B parameters, 24 attention heads, 4 KV heads, head dim 256
  • —Vocab: 248,320 tokens, native context: 262,144 tokens

Quantization Pipeline

Built with super-quant:

  1. 1.Convert HF to F16 GGUF (with MTP tensors) + mmproj GGUF (vision)
  2. 2.Multi-domain calibration data from 13 HF datasets, special tokens stripped
  3. 3.GPU-native importance matrix generation (PyTorch, 65k context) + weighted merge
  4. 4.Per-tensor sensitivity analysis (KL divergence probing against F16 logits)
  5. 5.Hybrid override generation — SSM at source precision, sensitivity-driven for the rest
  6. 6.Quantize with per-tensor overrides + imatrix
  7. 7.Benchmark: throughput + perplexity + KL divergence vs F16

Key Differences from Previous Release

AspectPreviousCurrent
CalibrationHad `<\endoftext\>` token leak (2,628 occurrences)Special tokens stripped automatically
Imatrix context32,76865,536
Imatrix generatorllama-imatrix (slow — PCIe D2H per tensor per chunk)PyTorch GPU-native (zero D2H during generation)
Sensitivity analysisllama.cpp subprocesses, ~600 GB disk I/O per runPyTorch in-place weight perturbation, zero disk I/O
SSM alpha/betaF32 (beyond source precision)F16 (matches source BF16)
SSM outQ8_0 (quantized)F16 (source precision preserved)
Attention tensorsMixed (f16/q80/q6k)All F16 (sensitivity-confirmed)
Norms/biasesImplicit (llama-quantize internal rules)Explicit F16 overrides
Total overrides339685

Links

Credits


License: Apache-2.0