scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants
21.9k
ORIGINAL MODEL BY LOGIC65, QUANTIZED FROM Whittle MoE 27B (A18B) v2.1
Models available here (go to logic65's official repo for Q4KM and Q8_0):
(note i have not tested any quantizations that are Q4 or above, let me know if one of the files don't work)
MODEL CARD FETCHED FROM: Whittle MoE 27B (A18B) v2.1 GGUF
Support this work
Every donation goes directly to GPU hours, and every GPU hour gets reported, including the failures: ko-fi.com/davida81328
Recommend using q8 for less looping and higher quality
Quantized builds of Whittle MoE 27B (A18B), the post hoc mixture of experts carved from Qwen3.8-27B and taught when to stop talking. v2.1 passed every release bar: 8 percent single turn loop rate (from 69 at first release), 22 percent structured (from 75), zero truncated answers, knowledge battery 28 of 39. All measurements, method, and failure history live on the main model card.
llama-server -m Whittle-MoE-27B-A18B-v2.1-Q4_K_M.gguf -ngl 99 -c 8192 -fa on --jinjaAny recent llama.cpp build with Qwen3.5 MoE support works. No fork, no patches.
