Uzbekswe/Muse-Glimmer-30B-Q4_K_M-GGUF
Muse Glimmer 30B — independent Q4KM GGUF
This repository contains an independently produced Q4_K_M GGUF conversion of Meta's Muse Glimmer 30B model. It is a research-preview artifact from the Quantized Muse Glimmer Research Lab.
This is not an official Meta quantization and is not byte-identical to the files in Meta's official Muse Glimmer GGUF release. It was created from the BF16 checkpoint with llama.cpp build b10353 using the Q4_K_M recipe.
Artifact
The source checkpoint used by the historical VESSL run was configured as mutable Hugging Face main; its exact historical commit was not exported. Future reproductions in the linked GitHub repository use an immutable source revision. See artifact-manifest.json for the complete provenance record.
Usage with llama.cpp
Use a recent llama.cpp build with Muse Glimmer support and enable the model's Jinja chat template:
llama-cli \
-m Muse-Glimmer-30B-custom-Q4_K_M.gguf \
--jinja \
--conversationFor exact runtime behavior, follow the upstream Muse Glimmer documentation and the linked research repository. The model is text-only in this release.
Preliminary measurements
The first VESSL tranche compared BF16, this custom Q4, and Meta's official text GGUF baselines. The results are directional and preliminary: speed and perplexity records were recovered as imported first-run measurements, while peak memory and billing snapshots were not available. The full report and machine-readable evidence are in the GitHub repository.
The custom Q4 measured 9.5372 perplexity versus 9.3473 for BF16 on the small held-out text sample. It measured 52.95 decode tokens/second versus 28.62 for BF16 in the imported first-run speed record. These numbers are not a broad quality guarantee.
Components not included
This repository contains only the main text GGUF. The mmproj vision projector and DFlash speculative-decoding component are not bundled. Obtain compatible components from Meta's official release and evaluate them as separate runtime components.
License and attribution
The upstream Muse Glimmer release is Apache-2.0. This derived GGUF is distributed under Apache-2.0 with attribution to Meta Superintelligence Labs. See the upstream model card and LICENSE for terms, limitations, and responsible-use guidance.
