ThakiCloud/Qwen3.8-27B-zhTW-cjk-suppressed
Qwen3.8-27B — Traditional Chinese Output Script Suppression
A mask that reduces, at the weight level, Qwen3.8-27B mixing in simplified Chinese + Japanese shinjitai when answering in Traditional Chinese. Only the corresponding token rows of lm_head are replaced.
Measurement results (2026-09-01)
Paired measurement over 300 prompts, B200 x1 · vLLM 0.28.0 · max_num_seqs=256 · compilation on. The mask was applied as a logit_bias -100 proxy (reproduces the true-weight result within -0.03pp at 27B).
This mask works. At T=1.0, the contamination rate drops from 40.0% to 1.33% (McNemar p=0.0).
Temperature drives the base contamination rate by 5.0x — serve at low temperature.
Capability regression
⚠️ When the difference is smaller than the minimum detectable difference, that means "not detected," not "no regression."
Cost probe (generation-level)
Requested 6 expressions this language must be able to use, and checked whether they survived in the answer — base 6/6 · masked 6/6, zero losses caused by the mask.
⚠️ Reading these numbers
- The contamination-character criterion uses the same repertoire as the mask's token criterion. It is not an independent metric.
- 300 prompts × 1 run, and there is no repeated run — run-to-run spread was not measured.
- The capability axes (coding·english) structurally have weak power to detect regression, since they don't use masked characters to begin with. That's why the cost probe exists separately.
This repo has no weights
Only lm_head.weight changes, so rather than making a 55.6GB copy, it's better to download the original and apply the mask.
hf download Qwen/Qwen3.8-27B --local-dir ./Qwen3.8-27B
python apply_mask.py --model-dir ./Qwen3.8-27B --mask mask_zh-TW_t1.jsonTakes a few minutes on CPU, no GPU needed. The script verifies the logit margin before writing, and after writing it reloads and asserts that no unintended tensor changed.
The mask
Rule: if a Hanja character is outside the Big5 repertoire, cut it. Simplified-only characters and Japanese-shinjitai-only characters are both caught together.
Tokenizer probe (what the code judged)
No measurement was run, but what the mask cuts and what it keeps can be checked without a GPU. This is on the default mask.
- preserved 7/7 — 電腦 · 資訊 · 網路 · 繁體 · 臺灣 · 軟體 · 國學會
- suppressed 7/7 — 电脑 · 资讯 · 网络 · 简体 · 软件 · 実際 · 気持
Running build_masks_multi.py re-executes this probe as-is.
Why cutting simplified alone isn't enough
Simplified is not the only contamination source in Traditional-Chinese context. Japanese shinjitai (実 気 広 経 歩) is also not Traditional, and a simplified/traditional converter doesn't catch it — s2t("実") returns 実 unchanged.
A Big5-repertoire criterion catches both at once, because every glyph form that isn't Traditional falls outside Big5.
⛔ Don't use this mask for Cantonese
5 of 11 Cantonese-specific characters (嘅 喺 啲 哋 嘢) fall outside plain Big5, so this mask cuts out Cantonese itself. Cantonese uses a separate repo → `Qwen3.8-27B-yue-cjk-suppressed`
Method
Replaces the target rows of lm_head.weight with a large negative multiple of the mean hidden-state direction.
W_i := -alpha * mu_h / ||mu_h||^2 (alpha = 200)lm_head has no bias, so zeroing a row only makes its logit 0, not −inf. The moment every other candidate is negative, 0 becomes the argmax. mu_h is measured via a forward pass.
⚠️ The PROBE_TEXTS in apply_mask.py are Korean sentences. If you're targeting this language, it's more correct to swap in sentences of that language — mu_h only means "this model's mean hidden state when using that language" if it's built that way. Swapping it changes ||mu_h||² too, so be sure to check the margin verification the script prints.
Limitations
- ⚠️ Serve at low `temperature`. In the Korean measurement, temperature drove the contamination rate by 5x (
T=1.09.33% vsT=0.01.92%). - The floor of pruning is not 0. Rare characters have no standalone token in the vocabulary and are assembled from byte fragments, and those byte tokens cannot be cut.
- English code-mixing cannot be fixed by this technique. English tokens must be kept alive for code, proper nouns, and units, so they are a contamination source that can't be cut.
- This is sanitization, not style improvement.
Prior work
`dnotitia/smoothie-qwen` distributes pre-adjusted checkpoints in the same direction (adjusting Qwen's lm_head to suppress Chinese). What this repo adds is a per-target-language repertoire criterion (legacy national encoding rather than a simplified/traditional axis) and a tier structure that differs by language.
Related: SASFT (ICLR 2026) · Korean token pruning · TLPO (ACL 2026)
License
apache-2.0 — same as base Qwen/Qwen3.8-27B. Distributes only the mask and scripts, and does not redistribute weights.
