ThakiCloud/Qwen3.8-27B-Human-KO-Safety-W4A16
Qwen3.8-27B-Human-KO-Safety-W4A16
This is the general-purpose quantized version for Hopper and above of `ThakiCloud/Qwen3.8-27B-Human-KO-Safety`. MLP weights are 4-bit (W4A16, grouped GPTQ), attention (selfattn·linearattn) is FP8 dynamic, and embeddings, lmhead, the recurrent input projection (`inproj_a/b`), and the vision tower are left as bf16. Disk size 20.6GiB (37% of bf16's 55.6GB).
What was re-measured for this version (2026-09-07, temperature 0, thinking off, same serving config as bf16)
In one line: abstention behavior is the same as bf16, and over-abstention and identity have each come down to the edge of their allowance. Capability gates (coding·KMMLU·instruction) were re-measured in the same run with bf16 as the reference arm and entered into the table above. The difference from the bf16 card's absolute values is because the runs differ; read the comparison only within the same run.
Quantization recipe
llm-compressor 0.13.0 GPTQ · calibration 1,024 samples × 2,048 tokens (chat format, 25% Korean) · dampening 0.01 · actorder static · scheme W4A16 (MLP) + FP8DYNAMIC (attention) · ignore: vision/visual, `lmhead, embedtokens`, `linearattn.inproja/b. Took 38 minutes (1x B200). Metadata is in quantizemeta.json`; the actual group and ignore lists are in `config.json`'s `quantizationconfig`.
Serving
Because attention is FP8, this requires a GPU with an FP8 kernel (SM89+: H100·H200·L40S·RTX 40/50). A100 (SM80) has not been verified and has no FP8 activation path, so correct operation is not guaranteed. On Blackwell, the NVFP4 version is faster.
vllm serve ThakiCloud/Qwen3.8-27B-Human-KO-Safety-W4A16 \
--max-model-len 32768 --kv-cache-dtype fp8 --enable-prefix-cachingvLLM ≥ 0.28 auto-detects the scheme from config.json (compressed-tensors).
Limitations
- The safety axis was measured with KoBBQ alone, and the table above is from a single build of this quantized checkpoint. Given the noise between requantization builds (our measured GSM8K figure is 3.56pp), a difference within 3pp is not attributed to the treatment.
- Identity's 94.6% differs from bf16's 96.2%. If self-introduction accuracy matters, use bf16.
- Other limitations and data provenance are the same as the bf16 card and
DATA_PROVENANCE.md.
License
Apache-2.0 (same as Qwen/Qwen3.8-27B). See LICENSE · NOTICE.
Paper
Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model — arXiv:2609.11291
The style alignment behind this line also moved two behaviors nobody trained for: abstention on ambiguous social questions (KoBBQ) and unprompted disclosure in securities guidance. Both moved through the emission policy — how often the model answers and how much it says — rather than through what it says when it does answer. Holding prompts, recipe, data volume and serving fixed and changing only the training target, three style seeds moved answer rate one way and three neutral seeds moved it the other (observed ranges do not overlap). Read this card's numbers with that in mind.
