tooltd/Qwen3.6-27B-mini-IQ4-XS-MTP-16GB-VRAM-GGUF
10194
Qwen3.6-27B Mini - IQ4_XS (GGUF)
An optimally sized quantized version of Qwen3.6-27B was created to fit 16GB with usable MTP.
Model Details
- Base Model: Qwen3.6-27B
- Quantization: IQ4_XS (Mini)
- BPW: 4.0012
- Quantized by: Using **Thireus' GGUF Tool Suite**
Quick test
- Wikitext-2-raw PPL: 7.0516 ± 0.04664
- Decoding speed RTX4070 Ti Super : 80 t/s
- Chessboard Test 😃

Hardware Requirements
- Fully fits on 16GB VRAM with 92K context using MTP + q4_0 KV cache.
- Can push even higher context by reducing KV cache further with TurboQuant or Kvarn.
Summary of tensor counts and bpw per qtype
QTYPE Count BPW Assigned GiB % Assigned Max GiB (all)
+f32 353 32 0.01 GiB - -
q8_0 6 8.5 0.00 GiB 0.01% 26.61
q6_K 101 6.5625 0.06 GiB 0.30% 20.55
q5_1 0 6 0.00 GiB 0.00% 18.78
q5_K 20 5.5 0.08 GiB 0.49% 17.22
q5_0 0 5.5 0.00 GiB 0.00% 17.22
q4_1 0 5 0.00 GiB 0.00% 15.65
q4_K 0 4.5 0.00 GiB 0.00% 14.09
q4_0 0 4.5 0.00 GiB 0.00% 14.09
iq4_nl 0 4.5 0.00 GiB 0.00% 14.09
iq4_xs 282 4.25 8.86 GiB 66.62% 13.31
q3_K 0 3.4375 0.00 GiB 0.00% 10.76
iq3_s 89 3.4375 3.51 GiB 32.59% 10.76
iq3_xxs 0 3.0625 0.00 GiB 0.00% 9.59
q2_K 0 2.625 0.00 GiB 0.00% 8.22
iq2_xs 0 2.3125 0.00 GiB 0.00% 7.24
iq2_xxs 0 2.0625 0.00 GiB 0.00% 6.46
iq1_m 0 1.75 0.00 GiB 0.00% 5.48
iq1_s 0 1.5625 0.00 GiB 0.00% 4.92
Average BPW: 4.0012