lambsea/Qwen3.6-27B-AEON-Ultimate-Uncensored-UD-GGUF
Qwen3.6-27B-AEON-Ultimate-Uncensored — GGUF (UD Quants)
Unsloth Dynamic-style (UD) GGUF quantizations of AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
Every quant uses per-tensor overrides (sensitivity-driven) + importance matrix (multi-domain calibration). All SSM recurrence tensors are preserved at source precision. MTP speculative decoding and vision (mmproj) are preserved.
Quant Comparison
Benchmarked on NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM), llama.cpp fork (a4501150/llama.cpp), pp=512, tg=128.
What Makes These Different
SSM Recurrence Preservation
Qwen3.6 is a hybrid GatedDeltaNet + attention model. 48 of 64 layers use a recurrent SSM where quantization error compounds across token positions. All SSM recurrence tensors are preserved at source precision (F16) — never quantized.
Per-Tensor Sensitivity Analysis
Each tensor group was probed by quantizing only that group to Q4_0 while keeping the rest at F16, then measuring KL divergence. The override generator assigns precision based on measured sensitivity:
685 total overrides — every non-FFN tensor has an explicit precision assignment. No dependence on llama-quantize's internal promotion rules.
Multi-Domain Calibration + GPU Imatrix
Calibrated on a balanced mix across 4 domains from 13 HF datasets:
Special tokens from source datasets are stripped automatically. Samples are kept whole — never truncated mid-conversation.
The importance matrix is generated with a PyTorch GPU-native generator (src/generate_imatrix.py) at 65,536 context — uses forward hooks to accumulate squared activations on GPU with zero PCIe D2H copies. Supports multi-GPU via device_map="auto".
Per-domain imatrices are merged with equal weights (DI-MATRIX approach).
MTP + Vision Preserved
- MTP (Multi-Token Prediction): Draft head (blk.64) pinned at F16. Use
--spec-type draft-mtp --spec-draft-n-max 3for ~1.5-2x faster generation. - Vision: mmproj file contains the full vision encoder. Use
--mmprojflag with llama-server for image/video understanding.
Files
Usage
llama-server (recommended)
# Q6_K with YaRN 512k context, 5 concurrent slots, MTP + vision
llama-server \
-m Qwen3.6-27B-AEON-UD-Q6_K.gguf \
--mmproj Qwen3.6-27B-AEON-mmproj-F16.gguf \
-ngl 99 \
--flash-attn \
-c 524288 \
--parallel 5 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
-kvu \
--cache-ram -1 \
--rope-scaling yarn \
--rope-scale 2.0 \
--yarn-orig-ctx 262144 \
--override-kv "qwen35.context_length=int:524288" \
--spec-type draft-mtp \
--spec-draft-n-max 3 \
--jinja \
--chat-template-kwargs '{"enable_thinking":true,"preserve_thinking":true}' \
--host 0.0.0.0 --port 8080Note: --spec-type draft-mtp requires llama.cpp b9375+. A custom fork adds DFlash speculative decoding and Blackwell-tuned flash attention.llama-cli
llama-cli \
-m Qwen3.6-27B-AEON-UD-Q6_K.gguf \
-ngl 99 \
--flash-attn \
-c 524288 \
--rope-scaling yarn \
--rope-scale 2.0 \
--yarn-orig-ctx 262144 \
--jinja \
--chat-template-kwargs '{"enable_thinking":true,"preserve_thinking":true}'Chat Template Notes
enable_thinkingactivates reasoning mode (chain-of-thought in<think>blocks)preserve_thinkingretains reasoning blocks in conversation history- No spaces after colons in the JSON — Qwen3.6's template parser is whitespace-sensitive
Architecture
Qwen3.6-27B is a hybrid SSM-attention model:
- 64 transformer layers + 1 MTP layer (blk.0-64)
- 48 SSM layers (GatedDeltaNet, no KV cache) + 16 full attention layers (every 4th layer)
- 27B parameters, 24 attention heads, 4 KV heads, head dim 256
- Vocab: 248,320 tokens, native context: 262,144 tokens
Quantization Pipeline
Built with super-quant:
- Convert HF to F16 GGUF (with MTP tensors) + mmproj GGUF (vision)
- Multi-domain calibration data from 13 HF datasets, special tokens stripped
- GPU-native importance matrix generation (PyTorch, 65k context) + weighted merge
- Per-tensor sensitivity analysis (KL divergence probing against F16 logits)
- Hybrid override generation — SSM at source precision, sensitivity-driven for the rest
- Quantize with per-tensor overrides + imatrix
- Benchmark: throughput + perplexity + KL divergence vs F16
Key Differences from Previous Release
Links
- Base model: AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
- Quantization pipeline: super-quant
- llama.cpp fork: a4501150/llama.cpp (DFlash, MTP fixes, Blackwell FA4)
Credits
- Base model: AEON-7
- Architecture: Qwen Team
- Quantization: llama.cpp
- Sensitivity methodology inspired by APEX quant research
- Calibration datasets: HuggingFaceH4, teknium, NousResearch, nvidia, open-r1, Salesforce, glaiveai, froggeric
License: Apache-2.0
