YuYu1015/Huihui-Qwen3.6-27B-abliterated-UD-Q4_K_XL-MTP-GGUF
Huihui-Qwen3.6-27B-abliterated-UD-Q4KXL-MTP-GGUF

English
Unsloth-style UD-Q4_K_XL quantization of huihui-ai/Huihui-Qwen3.6-27B-abliterated for llama.cpp, with the built-in MTP (Multi-Token Prediction) head fully preserved at Q8_0 for native speculative decoding.
Model Details
Per-Tensor Precision (UD Mask)
Notes on the Source Checkpoint
- huihui's abliterated checkpoint already contains the MTP head. Common wisdom says abliteration drops MTP, but
convert_hf_to_gguf.pypicks up all 4blk.64.nextn.*tensors directly — no Qwen-official MTP graft needed. - llama.cpp drops the vision tower automatically. The base model is
Qwen3_5ForConditionalGeneration(multimodal), but GGUF only packs the LLM portion. Text-only inference works out of the box.
Serving with llama.cpp (RTX 3090)
llama-server \
-m Huihui-Qwen3.6-27B-abliterated-UD-Q4_K_XL-MTP.gguf \
-fit off \
-c 65536 \
-np 1 \
-fa on \
-ngl 99 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--spec-type draft-mtp \
--spec-draft-n-max 2 \
--host 0.0.0.0 --port 8080Performance (RTX 3090 + Xeon E5-2696 v3)
Note on the CPU: the test rig uses an old Intel Xeon E5-2696 v3 (2.30 GHz, 2015). Modern CPUs (Zen 4 / Raptor Lake or newer) consistently push 30–50% higher tok/s on the same RTX 3090 because llama.cpp's hybrid GDN path still touches CPU per step. These numbers are a floor, not a ceiling.
Common Deployment Pitfalls
- Auto-fit abort on newer llama.cpp:
failed to fit params to free device memory, n_gpu_layers already set by user→ add-fit off. - Old llama.cpp can't load this:
missing tensor 'blk.64.ssm_conv1d.weight'→ upgrade to b9200+.
Safety Warning
This model has safety filtering removed (abliterated) and may generate sensitive, controversial, or inappropriate content. Users are solely responsible for all consequences arising from its use. Please ensure usage complies with local laws and ethical standards. Not suitable for public-facing or production applications.
Credits
- Original Model: Qwen/Qwen3.6-27B by Alibaba Qwen Team
- Abliteration: huihui-ai
- GGUF UD Quantization: YuYu1015
- UD Recipe Inspiration: Unsloth Dynamic Quants
繁體中文
huihui-ai/Huihui-Qwen3.6-27B-abliterated 的 Unsloth 風格 UD-Q4_K_XL 量化版本,針對 llama.cpp 部署,並完整保留內建 MTP(Multi-Token Prediction)head(Q8_0),原生支援投機解碼。
模型資訊
逐 Tensor 精度(UD 遮罩)
來源 Checkpoint 註記
- huihui abliterated checkpoint 已內含 MTP head。 主流說法是「abliteration 會弄丟 MTP」,但實際
convert_hf_to_gguf.py可直接抓到blk.64.nextn.*共 4 個 tensor — 不需嫁接 Qwen 官方 MTP。 - llama.cpp 自動跳過視覺塔。 基礎模型是
Qwen3_5ForConditionalGeneration(多模態),但 GGUF 只包 LLM 部分。純文字推理直接可用。
使用 llama.cpp 部署(RTX 3090)
llama-server \
-m Huihui-Qwen3.6-27B-abliterated-UD-Q4_K_XL-MTP.gguf \
-fit off \
-c 65536 \
-np 1 \
-fa on \
-ngl 99 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--spec-type draft-mtp \
--spec-draft-n-max 2 \
--host 0.0.0.0 --port 8080效能(RTX 3090 + Xeon E5-2696 v3)
CPU 說明: 測試機使用 2015 年的舊 Intel Xeon E5-2696 v3(2.30 GHz)。換成現代 CPU(Zen 4 / Raptor Lake 以上)在同樣 RTX 3090 上通常可再提升 30–50% tok/s,因為 llama.cpp 的 hybrid GDN 路徑每步仍需 CPU 介入。這個數字是下限,不是上限。
部署常見地雷
- 新版 llama.cpp auto-fit abort:
failed to fit params to free device memory, n_gpu_layers already set by user→ 加-fit off。 - 舊版 llama.cpp 載不了:
missing tensor 'blk.64.ssm_conv1d.weight'→ 升級到 b9200+。
安全警告
此模型已移除安全過濾機制(abliterated),可能產生敏感、爭議性或不當內容。使用者須自行承擔所有風險與法律責任,並確保使用方式符合當地法規與倫理標準。不適用於公開或生產環境。
致謝
- 原始模型:Qwen/Qwen3.6-27B,Alibaba Qwen 團隊
- 去審查:huihui-ai
- GGUF UD 量化:YuYu1015
- UD Recipe 靈感:Unsloth Dynamic Quants
