CoolFace
Modelpublic

RukaRat/Qwen3.8-27B-INT8-W8A8-imatrix-MTP

sourceHugging Faceapache-2.0updated 4d agoView on Hugging Face
8likes1.4kdownloads
14 commits on main
23f7b024d ago

re-upload model-00002-of-00002.safetensors from local Aug-14 build with xet disabled (2026-09-23)

RukaRat
ed1f2974d ago

re-upload model-00001-of-00002.safetensors from local Aug-14 build with xet disabled (2026-09-23)

RukaRat
cf3d6ad8d ago

Add the abliterated W8A8 build to the comparison; ground-truth scores for all abliterated builds

RukaRat
12ce6308d ago

Explain the SGLang single-card wall: MTP draft weights counted separately, 99.92%% of a 24GB card

RukaRat
88bfde08d ago

Correct SGLang single-card claim: it does not start (both failure modes documented); remove unreproducible 441k figure

RukaRat
e0168478d ago

Correct context table: plain does run MTP on one card (20k ctx), add SGLang columns, fix unsupported quality figure

RukaRat
b1ef1788d ago

Add 4-bit variant comparison and links to the W4A16 releases

RukaRat
aa9a64624d ago

Update card: SGLang is now the primary serving config; vLLM section corrected

RukaRat
d4680bb1mo ago

Trim RAM and tool-calling sections for readability

RukaRat
f707e611mo ago

Add YaRN context extension, KV cache notes, RAM requirement, tool-calling notes

RukaRat
b05485a1mo ago

Add INT8 W8A8 (imatrix-mse) weights, MTP and vision preserved

RukaRat
be22c841mo ago

Add LICENSE

RukaRat
9a398981mo ago

Add README.md

RukaRat
7ce03ea1mo ago

initial commit

RukaRat