sakamakismile/Huihui-Qwen3.5-27B-abliterated-NVFP4
Huihui-Qwen3.5-27B-abliterated-NVFP4
NVFP4 quantized version of huihui-ai/Huihui-Qwen3.5-27B-abliterated — an abliterated (uncensored) Qwen 3.5 27B dense model with multimodal capability and MTP (Multi-Token Prediction) support.
~52 GB → 20.6 GB with high-quality 512-sample calibration. Fits on a single NVIDIA Blackwell GPU.
Why This Model
- Uncensored — abliterated, no refusals for local agent workflows
- Deep reasoning — all responses start with structured "thinking process" chains
- 262K context — longest context window in its class
- MTP ready — Multi-Token Prediction head preserved in BF16 for speculative decoding
- Multimodal — vision tower preserved at full precision (BF16)
- Tool-call capable — works with vLLM
--enable-auto-tool-choice --tool-call-parser qwen3_xml
Key Specs
Quickstart
vLLM (recommended)
vllm serve Lna-Lab/Huihui-Qwen3.5-27B-abliterated-NVFP4 \
--max-model-len 32768 \
--reasoning-parser qwen3With tool calling
vllm serve Lna-Lab/Huihui-Qwen3.5-27B-abliterated-NVFP4 \
--max-model-len 32768 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--kv-cache-dtype fp8With MTP speculative decoding
vllm serve Lna-Lab/Huihui-Qwen3.5-27B-abliterated-NVFP4 \
--max-model-len 32768 \
--reasoning-parser qwen3 \
--speculative-config '{"method":"mtp","num_speculative_tokens":1}'Docker
docker run --gpus '"device=0"' -p 8016:8016 \
-v /path/to/model:/models/current:ro \
--shm-size 16gb \
-e VLLM_NVFP4_GEMM_BACKEND=marlin \
vllm/vllm-openai:cu130-nightly \
vllm serve /models/current --port 8016 --max-model-len 32768 \
--reasoning-parser qwen3Python
from vllm import LLM, SamplingParams
llm = LLM(
model="Lna-Lab/Huihui-Qwen3.5-27B-abliterated-NVFP4",
max_model_len=32768,
gpu_memory_utilization=0.90,
)
output = llm.generate(
["Implement a thread-safe LRU cache in Python with O(1) operations."],
SamplingParams(max_tokens=1024, temperature=0.3),
)
print(output[0].outputs[0].text)Benchmark
Tested on a single NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM).
Sustained throughput: ~59 tok/s (single GPU, post-warmup).
Note: This is a 27B dense model (all parameters active), so per-token speed is lower than MoE models like Gemma 4 26B-A4B (~130 tok/s with only 3.8B active). However, the reasoning depth per token is significantly higher.
Quantization Details
Recipe
recipe = QuantizationModifier(
targets=["Linear"],
ignore=["lm_head", "re:.*visual.*", "re:.*in_proj_a$", "re:.*in_proj_b$"],
scheme="NVFP4",
)Following the proven recipe from lyf/Qwen3.5-27B-Uncensored-HauhauCS-Aggressive-NVFP4.
What's quantized, what's not
- Quantized (NVFP4): All
Linearlayers in the text model - Kept in BF16:
lm_head, visual encoder, linear attention projections (in_proj_a,in_proj_b), MTP head
Calibration
- Dataset: neuralmagic/calibration (LLM split)
- Samples: 512 (high-quality calibration)
- Max sequence length: 4096
MTP (Multi-Token Prediction)
MTP tensors are grafted from the original BF16 checkpoint using save_mtp_tensors_to_checkpoint. This preserves the speculative decoding head at full precision, enabling ~3x speedup with num_speculative_tokens=1.
Reproduction
from compressed_tensors.utils import save_mtp_tensors_to_checkpoint
from transformers import Qwen3_5ForConditionalGeneration, AutoProcessor, AutoTokenizer
from datasets import load_dataset
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
import torch
MODEL_ID = "huihui-ai/Huihui-Qwen3.5-27B-abliterated"
OUTPUT = "Huihui-Qwen3.5-27B-abliterated-NVFP4"
model = Qwen3_5ForConditionalGeneration.from_pretrained(MODEL_ID, dtype="auto", trust_remote_code=True)
processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
recipe = QuantizationModifier(
targets=["Linear"],
ignore=["lm_head", "re:.*visual.*", "re:.*in_proj_a$", "re:.*in_proj_b$"],
scheme="NVFP4",
)
ds = load_dataset("neuralmagic/calibration", name="LLM", split="train[:512]")
def preprocess(example):
messages = [
{"role": m["role"], "content": [{"type": "text", "text": m["content"]}]}
for m in example["messages"]
]
return processor.apply_chat_template(
messages, return_tensors="pt", padding=False, truncation=True,
max_length=4096, tokenize=True, add_special_tokens=False,
return_dict=True, add_generation_prompt=False,
)
ds = ds.map(preprocess, batched=False, remove_columns=ds.column_names)
def data_collator(batch):
assert len(batch) == 1
return {
key: (torch.tensor(value) if key != "pixel_values"
else torch.tensor(value, dtype=torch.bfloat16).squeeze(0))
for key, value in batch[0].items()
}
oneshot(
model=model, recipe=recipe, dataset=ds,
max_seq_length=4096, num_calibration_samples=512,
data_collator=data_collator,
)
model.save_pretrained(OUTPUT, save_compressed=True)
processor.save_pretrained(OUTPUT)
tokenizer.save_pretrained(OUTPUT)
save_mtp_tensors_to_checkpoint(source_model=MODEL_ID, dest_dir=OUTPUT)Environment
Requirements
- GPU: NVIDIA Blackwell (RTX 5090, RTX PRO 6000, B200, etc.) — NVFP4 requires SM 120
- VRAM: ~21 GB minimum (model only), ~90 GB for 262K context
- Software: vLLM nightly (cu130 build)
Notes
- This is an abliterated (uncensored) model. Use responsibly.
- Vision tower is kept in BF16 — multimodal capabilities are preserved.
- MTP head is kept in BF16 — speculative decoding works out of the box.
- NVFP4 is a Blackwell-specific format. This will not work on Ampere/Hopper GPUs.
- For maximum context length (262K), use
--kv-cache-dtype fp8to fit in 96 GB.
Credits
- Base model: huihui-ai (abliteration)
- Original model: Qwen (Qwen 3.5)
- Quantization recipe: lyf/HauhauCS (proven path)
- Quantization tool: vllm-project/llm-compressor
Support the Base Model Author
If you find this model useful, please consider supporting huihui-ai:
- Ko-fi: https://ko-fi.com/huihuiai
- Bitcoin:
bc1qqnkhuchxw0zqjh2ku3lu4hq45hc6gy84uk70ge
