zrlu/Huihui-Qwen3.8-27B-abliterated-GPTQ-Int4-sym-G128-MTP-BF16-B70
Huihui-Qwen3.8-27B-abliterated-GPTQ-Int4-sym-G128-MTP-BF16 (B70)
GPTQ-INT4 quantization, purpose-built and tested for single Intel Arc Pro B70 (Xe2, Battlemage) inference via vLLM XPU on Windows Docker Desktop + WSL2 — and for pi/opencode agent workloads on the same card.
- Model name: Huihui-Qwen3.8-27B-abliterated-GPTQ-Int4-sym-G128-MTP-BF16-B70
- Original model: huihui-ai/Huihui-Qwen3.8-27B-abliterated (abliterated update: layers 18–51 ablated only; MTP and vision untouched), itself a derivative of Qwen/Qwen3.8-27B (Apache-2.0).
- Type / purpose-built: quantized derivative checkpoint (GPTQ-INT4), MTP (spec-decode) heads preserved in BF16 — the draft INT4 quantization overlay is disabled on this model variant (
DRAFT_INT4=0); see the deploy repository's README ("Runtime patches") for the shape mismatch behind that.
License
Apache-2.0 — inherited unchanged from the base model. See LICENSE and the base model card. This is a weight-only quantization; no rights are widened.
Quantization (self-hosted, reproducible)
- Tool:
gptqmodel==7.3.2; run on a CUDA GPU (RTX 5090) with lazy checkpoint loading; calibration: wikitext-2 (or bundled fallback text). - Contract:
bits=4, group_size=128, desc_act=false, sym=true, lm_head=false, dynamic={"-:.*mtp.*":{}}(MTP tensors stay BF16). - Input 51.75 GB (BF16) → output 18.22 GB (−64.8%), 5 safetensors shards, config arch
Qwen3_5ForConditionalGeneration,image_token_id 248056. - 400 quantized INT4 weight tensors + 15 preserved BF16
mtp.*tensors.
Serving (Arc Pro B70)
Served with vLLM XPU 0.28.0 (kernels 0.1.12.3), MTP3 speculative decoding with BF16 draft (DRAFT_INT4=0), prefix caching, qwen3_xml tool-call parser, fp8 KV, KV cache 4.3 GiB (KV_CACHE_MEMORY_BYTES=4617089843), max-num-seqs=1. No repetition/presence penalty — sampling is pure Qwen defaults from the model's generation_config.json (temperature 1.0 / topk 20 / topp 0.95); the deploy repo's patch_tile_mask.py boot patch covers the vLLM 0.28.0 !-degeneration bug instead. Benchmark/report details live in the deployment repository (benchmarks/ section).
One-click container: zrlu/qwen38-27b-arc-pro-b70:latest on Docker Hub (entrypoint auto-downloads this repo to /model on first start; switch to any other published B70 quant with -e HF_REPO=zrlu/<repo>), Dockerfile and pi-agent setup in the deploy repository.
Usage
# standalone / eval
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("zrlu/Huihui-Qwen3.8-27B-abliterated-GPTQ-Int4-sym-G128-MTP-BF16-B70")
m = AutoModelForCausalLM.from_pretrained(
"zrlu/Huihui-Qwen3.8-27B-abliterated-GPTQ-Int4-sym-G128-MTP-BF16-B70",
device_map="auto")Signed: zrlu. GPU: Arc Pro B70 (0xe223) + RTX 5090 (quantization only).
