CoolFace
Modelpublic

sakamakismile/Huihui-Qwen3.5-27B-abliterated-NVFP4

sourceHugging Faceotherupdated 6mo agoView on Hugging Face
1likes101downloads
Model Card

Huihui-Qwen3.5-27B-abliterated-NVFP4

NVFP4 quantized version of huihui-ai/Huihui-Qwen3.5-27B-abliterated — an abliterated (uncensored) Qwen 3.5 27B dense model with multimodal capability and MTP (Multi-Token Prediction) support.

~52 GB → 20.6 GB with high-quality 512-sample calibration. Fits on a single NVIDIA Blackwell GPU.

Why This Model

  • —Uncensored — abliterated, no refusals for local agent workflows
  • —Deep reasoning — all responses start with structured "thinking process" chains
  • —262K context — longest context window in its class
  • —MTP ready — Multi-Token Prediction head preserved in BF16 for speculative decoding
  • —Multimodal — vision tower preserved at full precision (BF16)
  • —Tool-call capable — works with vLLM --enable-auto-tool-choice --tool-call-parser qwen3_xml

Key Specs

Base modelhuihui-ai/Huihui-Qwen3.5-27B-abliterated
ArchitectureQwen 3.5 Dense — 27B parameters, 64 layers
QuantizationNVFP4 W4A4 (weights FP4, activations FP4, scales FP8)
Formatcompressed-tensors (native vLLM support)
Toolvllm-project/llm-compressor (main)
Calibration512 samples, neuralmagic/calibration, seq_len=4096
Size20.6 GB
Max context262,144 tokens
RequiresNVIDIA Blackwell GPU (SM 120), vLLM nightly (cu130)

Quickstart

vLLM (recommended)

bash
vllm serve Lna-Lab/Huihui-Qwen3.5-27B-abliterated-NVFP4 \
    --max-model-len 32768 \
    --reasoning-parser qwen3

With tool calling

bash
vllm serve Lna-Lab/Huihui-Qwen3.5-27B-abliterated-NVFP4 \
    --max-model-len 32768 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_xml \
    --kv-cache-dtype fp8

With MTP speculative decoding

bash
vllm serve Lna-Lab/Huihui-Qwen3.5-27B-abliterated-NVFP4 \
    --max-model-len 32768 \
    --reasoning-parser qwen3 \
    --speculative-config '{"method":"mtp","num_speculative_tokens":1}'

Docker

bash
docker run --gpus '"device=0"' -p 8016:8016 \
    -v /path/to/model:/models/current:ro \
    --shm-size 16gb \
    -e VLLM_NVFP4_GEMM_BACKEND=marlin \
    vllm/vllm-openai:cu130-nightly \
    vllm serve /models/current --port 8016 --max-model-len 32768 \
    --reasoning-parser qwen3

Python

python
from vllm import LLM, SamplingParams

llm = LLM(
    model="Lna-Lab/Huihui-Qwen3.5-27B-abliterated-NVFP4",
    max_model_len=32768,
    gpu_memory_utilization=0.90,
)

output = llm.generate(
    ["Implement a thread-safe LRU cache in Python with O(1) operations."],
    SamplingParams(max_tokens=1024, temperature=0.3),
)
print(output[0].outputs[0].text)

Benchmark

Tested on a single NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM).

TestSpeedTokensResult
English (system design)59.2 tok/s512PASS
Code (async scheduler)59.2 tok/s512PASS
Math (Bayes' theorem)59.2 tok/s512PASS
Japanese (technical writing)59.0 tok/s512PASS

Sustained throughput: ~59 tok/s (single GPU, post-warmup).

Note: This is a 27B dense model (all parameters active), so per-token speed is lower than MoE models like Gemma 4 26B-A4B (~130 tok/s with only 3.8B active). However, the reasoning depth per token is significantly higher.

Quantization Details

Recipe

python
recipe = QuantizationModifier(
    targets=["Linear"],
    ignore=["lm_head", "re:.*visual.*", "re:.*in_proj_a$", "re:.*in_proj_b$"],
    scheme="NVFP4",
)

Following the proven recipe from lyf/Qwen3.5-27B-Uncensored-HauhauCS-Aggressive-NVFP4.

What's quantized, what's not

  • —Quantized (NVFP4): All Linear layers in the text model
  • —Kept in BF16: lm_head, visual encoder, linear attention projections (in_proj_a, in_proj_b), MTP head

Calibration

  • —Dataset: neuralmagic/calibration (LLM split)
  • —Samples: 512 (high-quality calibration)
  • —Max sequence length: 4096

MTP (Multi-Token Prediction)

MTP tensors are grafted from the original BF16 checkpoint using save_mtp_tensors_to_checkpoint. This preserves the speculative decoding head at full precision, enabling ~3x speedup with num_speculative_tokens=1.

Reproduction

python
from compressed_tensors.utils import save_mtp_tensors_to_checkpoint
from transformers import Qwen3_5ForConditionalGeneration, AutoProcessor, AutoTokenizer
from datasets import load_dataset
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
import torch

MODEL_ID = "huihui-ai/Huihui-Qwen3.5-27B-abliterated"
OUTPUT = "Huihui-Qwen3.5-27B-abliterated-NVFP4"

model = Qwen3_5ForConditionalGeneration.from_pretrained(MODEL_ID, dtype="auto", trust_remote_code=True)
processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)

recipe = QuantizationModifier(
    targets=["Linear"],
    ignore=["lm_head", "re:.*visual.*", "re:.*in_proj_a$", "re:.*in_proj_b$"],
    scheme="NVFP4",
)

ds = load_dataset("neuralmagic/calibration", name="LLM", split="train[:512]")

def preprocess(example):
    messages = [
        {"role": m["role"], "content": [{"type": "text", "text": m["content"]}]}
        for m in example["messages"]
    ]
    return processor.apply_chat_template(
        messages, return_tensors="pt", padding=False, truncation=True,
        max_length=4096, tokenize=True, add_special_tokens=False,
        return_dict=True, add_generation_prompt=False,
    )

ds = ds.map(preprocess, batched=False, remove_columns=ds.column_names)

def data_collator(batch):
    assert len(batch) == 1
    return {
        key: (torch.tensor(value) if key != "pixel_values"
              else torch.tensor(value, dtype=torch.bfloat16).squeeze(0))
        for key, value in batch[0].items()
    }

oneshot(
    model=model, recipe=recipe, dataset=ds,
    max_seq_length=4096, num_calibration_samples=512,
    data_collator=data_collator,
)

model.save_pretrained(OUTPUT, save_compressed=True)
processor.save_pretrained(OUTPUT)
tokenizer.save_pretrained(OUTPUT)
save_mtp_tensors_to_checkpoint(source_model=MODEL_ID, dest_dir=OUTPUT)

Environment

PackageVersion
torch2.11.0+cu130
transformers5.5.4
llmcompressor0.1.dev (main @ 3084520)
compressed-tensors0.15.1a20260414
CUDA13.0

Requirements

  • —GPU: NVIDIA Blackwell (RTX 5090, RTX PRO 6000, B200, etc.) — NVFP4 requires SM 120
  • —VRAM: ~21 GB minimum (model only), ~90 GB for 262K context
  • —Software: vLLM nightly (cu130 build)

Notes

  • —This is an abliterated (uncensored) model. Use responsibly.
  • —Vision tower is kept in BF16 — multimodal capabilities are preserved.
  • —MTP head is kept in BF16 — speculative decoding works out of the box.
  • —NVFP4 is a Blackwell-specific format. This will not work on Ampere/Hopper GPUs.
  • —For maximum context length (262K), use --kv-cache-dtype fp8 to fit in 96 GB.

Credits

Support the Base Model Author

If you find this model useful, please consider supporting huihui-ai: