RukaRat/Qwen3.8-27B-INT8-W8A8-imatrix-MTP
re-upload model-00002-of-00002.safetensors from local Aug-14 build with xet disabled (2026-09-23)
re-upload model-00001-of-00002.safetensors from local Aug-14 build with xet disabled (2026-09-23)
Add the abliterated W8A8 build to the comparison; ground-truth scores for all abliterated builds
Explain the SGLang single-card wall: MTP draft weights counted separately, 99.92%% of a 24GB card
Correct SGLang single-card claim: it does not start (both failure modes documented); remove unreproducible 441k figure
Correct context table: plain does run MTP on one card (20k ctx), add SGLang columns, fix unsupported quality figure
Add 4-bit variant comparison and links to the W4A16 releases
Update card: SGLang is now the primary serving config; vLLM section corrected
Trim RAM and tool-calling sections for readability
Add YaRN context extension, KV cache notes, RAM requirement, tool-calling notes
Add INT8 W8A8 (imatrix-mse) weights, MTP and vision preserved
Add LICENSE
Add README.md
initial commit
