TheUnderscore/Swift-Qwen3.8-27b-W4A16-AWQ
133.8k
Swift-Qwen3.8-27B-W4A16-AWQ
W4A16 (4-bit weights, 16-bit activations) AWQ compressed-tensors quantization of `ukisai/Swift-Qwen3.8-27b`.
Quantization method
- Scheme:
W4A16_ASYM— 4-bit asymmetric per-group quantization (group size 128) of allLinearweights, stored in the compressed-tensors pack-quantized format (weight_packed/weight_scale/weight_zero_point/weight_shape), which LMDeployturbomindauto-detects and loads natively (including the MTP head and vision tower, which stay BF16). - Tooling: llmcompressor one-shot offline quantization with CPU offloading (
compressed_tensors.offload.load_offloaded_model), so the full-precision source fits on a 2×16 GB VRAM setup. - AWQ activation smoothing:
AWQModifierwith the layer-scoped hybrid-attention mappings frombuild_hybrid_attention_mappings— full-attentioninput_layernorm→self_attn.q/k/v,post_attention_layernorm→mlp.gate/up, andmlp.up_proj→mlp.down_proj, withduo_scaling="both"and CPU offload, followed by W4A16 quantization. This layer-scoped recipe is required for hybrid-attention (Qwen3.5-family) architectures — grouped-regex smoothing or mismatched mappings corrupt decoding. - Unquantized (kept BF16): embeddings,
lm_head, norms,linear_attn.in_proj_a/b, the vision tower, and the MTP head. - Run command:
CUDA_VISIBLE_DEVICES=1,2 python3 quantize-awq-hybrid.py \
--model_path ./Swift-Qwen3.8-27b \
--quant_path ./Swift-Qwen3.8-27b-W4A16-AWQ \
--offload_dir ./Swift-Qwen3.8-27b-W4A16-AWQ(Default calibration: UltraChat 200k train_sft.)
Recommended parameters
From the base model: temperature 1.0, topp 0.95, topk 20, minp 0, presencepenalty 0, repetition_penalty 1.0.
Usage
Tested with LMDeploy turbomind:
from lmdeploy import pipeline, TurbomindEngineConfig
pipe = pipeline(
"TheUnderscore/Swift-Qwen3.8-27b-W4A16-AWQ",
backend_config=TurbomindEngineConfig(
tp=2,
model_format="compressed-tensors",
language_model_only=True,
),
)
print(pipe("Hello, who are you?").text)License and access
Swift weights are distributed through gated access under the Swift Open License v1.0; this quantization inherits that license from the base model. Personal, research, educational, evaluation, and commercial use are free for individuals and organizations with annual recurring revenue, including affiliates, of up to US$1,000,000. Above that threshold, commercial use requires a separate Swift Enterprise License. Contact UkisAI for terms.
Files
quantize-awq-hybrid.py— the script used to produce this quantization (CPU-offloaded, DDP/torchrun, produces properly numbered-of-Nshards).--offload_dirselects where per-rank CPU offload temp folders live (defaults to the current working directory).model-nonquant.safetensors— unquantized tensors (mtp.*andmodel.visual.*) preserved BF16 so the full model architecture is loadable.
