CoolFace
Modelpublic

scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
2likes1.9kdownloads
Model Card

ORIGINAL MODEL BY LOGIC65, QUANTIZED FROM Whittle MoE 27B (A18B) v2.1

Models available here (go to logic65's official repo for Q4KM and Q8_0):

filesizeuse
Whittle-MoE-27B-A18B-v2.1-Q2_K.gguf11.3 GBsmallest usable file
Whittle-MoE-27B-A18B-v2.1-Q3_K_S.gguf12.7 GB
Whittle-MoE-27B-A18B-v2.1-Q3_K_M.gguf13.9 GBlikely best quality possible for 16GB VRAM, if not Q3KS
Whittle-MoE-27B-A18B-v2.1-Q3_K_L.gguf14.7 GB
Whittle-MoE-27B-A18B-v2.1-Q4_K_S.gguf16.1 GB
`Whittle-MoE-27B-A18B-v2.1-Q4_K_M.gguf`17.4 GBrecommended; fits 24 GB VRAM, splits across two 12 GB cards
Whittle-MoE-27B-A18B-v2.1-Q5_K_S.gguf19.0 GBtry using if spare headroom is available
Whittle-MoE-27B-A18B-v2.1-Q5_K_M.gguf19.9 GB
Whittle-MoE-27B-A18B-v2.1-Q6_K.gguf23.1 GB
`Whittle-MoE-27B-A18B-v2.1-Q8_0.gguf`28.7 GBhigher quality reference
Whittle-MoE-27B-A18B-v2.1-F16.gguf53.9 GBuse for creating future quantizations

(note i have not tested any quantizations that are Q4 or above, let me know if one of the files don't work)


MODEL CARD FETCHED FROM: Whittle MoE 27B (A18B) v2.1 GGUF

Support this work

Every donation goes directly to GPU hours, and every GPU hour gets reported, including the failures: ko-fi.com/davida81328

Recommend using q8 for less looping and higher quality

Quantized builds of Whittle MoE 27B (A18B), the post hoc mixture of experts carved from Qwen3.8-27B and taught when to stop talking. v2.1 passed every release bar: 8 percent single turn loop rate (from 69 at first release), 22 percent structured (from 75), zero truncated answers, knowledge battery 28 of 39. All measurements, method, and failure history live on the main model card.

filesizeuse
`Whittle-MoE-27B-A18B-v2.1-Q4_K_M.gguf`17.4 GBrecommended; fits 24 GB VRAM, splits across two 12 GB cards
`Whittle-MoE-27B-A18B-v2.1-Q8_0.gguf`28.7 GBhigher quality reference
bash
llama-server -m Whittle-MoE-27B-A18B-v2.1-Q4_K_M.gguf -ngl 99 -c 8192 -fa on --jinja

Any recent llama.cpp build with Qwen3.5 MoE support works. No fork, no patches.