CoolFace
Modelpublic

AIconjured/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-MTP-GGUF-NVFP4

sourceHugging Faceotherupdated 22d agoView on Hugging Face
5likes3.2kdownloads
Model Card

Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-MTP-GGUF-NVFP4

Uncensored Qwen3.5 9B (HauhauCS Aggressive) · NVFP4 quantized · MTP head grafted · vision-capable

A hand-built GGUF of HauhauCS's uncensored Qwen3.5-9B fine-tune, quantized with a custom NVFP4 recipe for Blackwell GPUs, with the Multi-Token-Prediction (MTP) head grafted in from the upstream base model, plus the original vision projector (mmproj) for image input.

Credits

This model is built on the work of two teams:

  • —The Qwen team (QwenLM) — built the base Qwen3.5-9B model, including its MTP head.
  • —The HauhauCS team (HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive) — tuned the base model into this uncensored "Aggressive" variant (0/465 refusals, fully unlocked with no capability loss) and provided the vision projector.

This repository (AIconjured) is the NVFP4 quantization, MTP graft, and packaging of that model.

Files

FileSizeDescription
Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-MTP-GGUF-NVFP4.gguf5.9 GBLLM: 442 tensors (427 main + 15 MTP), block_count=33, nextn_predict_layers=1
mmproj-Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-BF16.gguf0.9 GBVision projector, clip arch, 334 tensors, BF16 (unchanged from upstream)

Size: ~5.9 GB (LLM) + ~0.9 GB (mmproj) — down from 18.8 GB in BF16.

Model facts

  • —Architecture: qwen35 — 32 trunk layers (hybrid DeltaNet SSM + full attention, interval 4), 9.2B params, 262K native context, 4096-dim embeddings, 16 heads / 4 KV heads
  • —Base model: HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive (BF16, 17.9 GB)
  • —MTP source: unsloth Qwen3.5-9B-MTP-GGUF (Q8_0 carrier)

Quantization recipe

The quantization is a mixed-precision recipe, not a single uniform type. It was built with llama.cpp's --tensor-type-file (regex → type) plus an imatrix generated from ~144 KB of mixed calibration text (PPL ≈ 5.18):

TypeTensorsWhatWhy
NVFP4192ffn_gate/up/down, attn_qkv, attn_gate, ssm_out, full-attn attn_q/k/v/outputAll the big GEMMs — where the bits are. NVFP4 (FP4 weights + FP8/UE4M3 scales) is the native Blackwell path (BLACKWELL_NATIVE_FP4), fastest on RTX 50-series.
Q8_066token_embd, output (LM head), attn_k, ssm_alpha, ssm_beta, ssm_conv1dSmall but accuracy-sensitive tensors; Q8_0 is near-lossless.
F32184All norm weights, ssm_a, ssm_dtNorms/decay constants are tiny but quantizing them causes noticeable quality loss; kept at full precision.

Recipe (llama.cpp --tensor-type-file format):

token_embd\.weight=q8_0
^output\.weight=q8_0
blk\.\d+\.attn_k\.weight=q8_0
blk\.\d+\.ssm_beta\.weight=q8_0
blk\.\d+\.ssm_alpha\.weight=q8_0
blk\.\d+\.ssm_conv1d\.weight=q8_0
blk\.\d+\.attn_output\.weight=nvfp4
blk\.\d+\.ffn_gate\.weight=nvfp4
blk\.\d+\.ffn_up\.weight=nvfp4
blk\.\d+\.ffn_down\.weight=nvfp4
blk\.\d+\.attn_qkv\.weight=nvfp4
blk\.\d+\.attn_gate\.weight=nvfp4
blk\.\d+\.ssm_out\.weight=nvfp4
blk\.\d+\.attn_q\.weight=nvfp4
blk\.\d+\.attn_v\.weight=nvfp4

MTP (Multi-Token Prediction)

The HauhauCS fine-tune ships without MTP tensors. The MTP head (15 tensors, blk.32.nextn.*) was grafted from unsloth's Q80 base-model carrier, and the metadata was set to `blockcount=33 + nextnpredictlayers=1` so the loader sees 32 trunk layers + 1 MTP layer at index 32.

Important: the MTP head comes from the base model, not the fine-tune. Speculative decoding is self-verifying (the main model accepts/rejects every draft token), so quality is guaranteed by the uncensored fine-tune — but the accept rate will be lower than a fine-tune-native MTP head, since the base head doesn't perfectly match the fine-tuned distribution.

To actually get the speedup you need an engine that implements MTP spec decode:

bash
llama-server -m Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-MTP-GGUF-NVFP4.gguf \
  --mmproj mmproj-Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-BF16.gguf \
  --spec-type draft-mtp -c 8192 -ngl 99

Ollama note: Ollama (tested 0.33.3) loads the MTP tensors but does not implement MTP speculative decoding — the tensors are ignored and the model runs as a plain 32-layer model. Text and vision both work fine; you just don't get the draft speedup.

Vision

Vision is a separate `clip`-arch mmproj GGUF (334 tensors, BF16, unmodified from upstream). The LLM file intentionally contains no vision tensors — the qwen35 arch loader rejects them, and this is the structure Ollama expects. Load both files:

FROM ./Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-MTP-GGUF-NVFP4.gguf
FROM ./mmproj-Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-BF16.gguf

The projector uses the qwen3vl_merger with a vision encoder; verified working end-to-end in Ollama (text + image).

Hardware / performance notes

  • —Built and verified on 2× RTX 5060 Ti (16 GB each, Blackwell, sm_120a) with a CUDA 12.8 llama.cpp build.
  • —All 32 layers + output offload to GPU; the model fits entirely in VRAM on a single 16 GB card at moderate context.
  • —NVFP4 kernels use the native Blackwell FP4 path; on non-Blackwell GPUs the tensors still load and run (emulated/dequantized), with less of a speed advantage.
  • —262K native context is declared; practical usable context is limited by VRAM.

Reproduction

  • —llama.cpp: 378aa2ebc (2026-08-27), CUDA build
  • —Quantize: llama-quantize --imatrix imatrix-9b.gguf --tensor-type-file recipe.txt ...
  • —Imatrix: 56 chunks × 2048 tokens, mixed wiki/prose calibration, final PPL 5.1784 ± 0.096
  • —MTP graft: tensor copy of the 15 blk.32.nextn.* tensors + qwen35.nextn_predict_layers=1, qwen35.block_count=33

License

This quantization is a derivative of HauhauCS's uncensored fine-tune of the Qwen3.5-9B base model. See the base model's license at HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive and the Qwen base model's license at QwenLM.