CoolFace
Modelpublic

zrlu/Huihui-Qwen3.8-27B-abliterated-GPTQ-Int4-sym-G128-MTP-BF16-B70

sourceHugging Faceapache-2.0updated 28d agoView on Hugging Face
2likes353downloads
Model Card

Huihui-Qwen3.8-27B-abliterated-GPTQ-Int4-sym-G128-MTP-BF16 (B70)

GPTQ-INT4 quantization, purpose-built and tested for single Intel Arc Pro B70 (Xe2, Battlemage) inference via vLLM XPU on Windows Docker Desktop + WSL2 — and for pi/opencode agent workloads on the same card.

  • —Model name: Huihui-Qwen3.8-27B-abliterated-GPTQ-Int4-sym-G128-MTP-BF16-B70
  • —Original model: huihui-ai/Huihui-Qwen3.8-27B-abliterated (abliterated update: layers 18–51 ablated only; MTP and vision untouched), itself a derivative of Qwen/Qwen3.8-27B (Apache-2.0).
  • —Type / purpose-built: quantized derivative checkpoint (GPTQ-INT4), MTP (spec-decode) heads preserved in BF16 — the draft INT4 quantization overlay is disabled on this model variant (DRAFT_INT4=0); see the deploy repository's README ("Runtime patches") for the shape mismatch behind that.

License

Apache-2.0 — inherited unchanged from the base model. See LICENSE and the base model card. This is a weight-only quantization; no rights are widened.

Quantization (self-hosted, reproducible)

  • —Tool: gptqmodel==7.3.2; run on a CUDA GPU (RTX 5090) with lazy checkpoint loading; calibration: wikitext-2 (or bundled fallback text).
  • —Contract: bits=4, group_size=128, desc_act=false, sym=true, lm_head=false, dynamic={"-:.*mtp.*":{}} (MTP tensors stay BF16).
  • —Input 51.75 GB (BF16) → output 18.22 GB (−64.8%), 5 safetensors shards, config arch Qwen3_5ForConditionalGeneration, image_token_id 248056.
  • —400 quantized INT4 weight tensors + 15 preserved BF16 mtp.* tensors.

Serving (Arc Pro B70)

Served with vLLM XPU 0.28.0 (kernels 0.1.12.3), MTP3 speculative decoding with BF16 draft (DRAFT_INT4=0), prefix caching, qwen3_xml tool-call parser, fp8 KV, KV cache 4.3 GiB (KV_CACHE_MEMORY_BYTES=4617089843), max-num-seqs=1. No repetition/presence penalty — sampling is pure Qwen defaults from the model's generation_config.json (temperature 1.0 / topk 20 / topp 0.95); the deploy repo's patch_tile_mask.py boot patch covers the vLLM 0.28.0 !-degeneration bug instead. Benchmark/report details live in the deployment repository (benchmarks/ section).

One-click container: zrlu/qwen38-27b-arc-pro-b70:latest on Docker Hub (entrypoint auto-downloads this repo to /model on first start; switch to any other published B70 quant with -e HF_REPO=zrlu/<repo>), Dockerfile and pi-agent setup in the deploy repository.

Usage

python
# standalone / eval
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("zrlu/Huihui-Qwen3.8-27B-abliterated-GPTQ-Int4-sym-G128-MTP-BF16-B70")
m = AutoModelForCausalLM.from_pretrained(
    "zrlu/Huihui-Qwen3.8-27B-abliterated-GPTQ-Int4-sym-G128-MTP-BF16-B70",
    device_map="auto")

Signed: zrlu. GPU: Arc Pro B70 (0xe223) + RTX 5090 (quantization only).