CoolFace
Modelpublic

mgaruccio/Muse-Glimmer-30B-GPTQ-Int4-sym-G128

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes49downloads
Model Card

Muse Glimmer 30B GPTQ Int4 G128

GPTQ 4-bit / group-128 / symmetric (desc_act=false, lm_head=false) of Meta Muse Glimmer 30B. Apache 2.0.

This is the target for the vLLM-XPU + DFlash B70 recipe. Pair it with `Muse-Glimmer-30B-assistant-GPTQ-Int4-sym-G128`.

  • —Quantizer: GPTQModel 7.3.2
  • —Vision / embeddings / lm_head left unquantized; serve with --language-model-only
  • —Do not treat a from-scratch requant as bit-identical to these shards
bash
hf download mgaruccio/Muse-Glimmer-30B-GPTQ-Int4-sym-G128 --local-dir ./models/target