CoolFace
Modelpublic

Uzbekswe/Muse-Glimmer-30B-Q4_K_M-GGUF

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
2likes274downloads
Model Card

Muse Glimmer 30B — independent Q4KM GGUF

This repository contains an independently produced Q4_K_M GGUF conversion of Meta's Muse Glimmer 30B model. It is a research-preview artifact from the Quantized Muse Glimmer Research Lab.

This is not an official Meta quantization and is not byte-identical to the files in Meta's official Muse Glimmer GGUF release. It was created from the BF16 checkpoint with llama.cpp build b10353 using the Q4_K_M recipe.

Artifact

FileSizeSHA-256
Muse-Glimmer-30B-custom-Q4_K_M.gguf16,935,294,592 bytes (15.77 GiB)dcd100005262563bdcef650b2de0b41f285570a5dc2760af63d546c224b2e21d

The source checkpoint used by the historical VESSL run was configured as mutable Hugging Face main; its exact historical commit was not exported. Future reproductions in the linked GitHub repository use an immutable source revision. See artifact-manifest.json for the complete provenance record.

Usage with llama.cpp

Use a recent llama.cpp build with Muse Glimmer support and enable the model's Jinja chat template:

bash
llama-cli \
  -m Muse-Glimmer-30B-custom-Q4_K_M.gguf \
  --jinja \
  --conversation

For exact runtime behavior, follow the upstream Muse Glimmer documentation and the linked research repository. The model is text-only in this release.

Preliminary measurements

The first VESSL tranche compared BF16, this custom Q4, and Meta's official text GGUF baselines. The results are directional and preliminary: speed and perplexity records were recovered as imported first-run measurements, while peak memory and billing snapshots were not available. The full report and machine-readable evidence are in the GitHub repository.

The custom Q4 measured 9.5372 perplexity versus 9.3473 for BF16 on the small held-out text sample. It measured 52.95 decode tokens/second versus 28.62 for BF16 in the imported first-run speed record. These numbers are not a broad quality guarantee.

Components not included

This repository contains only the main text GGUF. The mmproj vision projector and DFlash speculative-decoding component are not bundled. Obtain compatible components from Meta's official release and evaluate them as separate runtime components.

License and attribution

The upstream Muse Glimmer release is Apache-2.0. This derived GGUF is distributed under Apache-2.0 with attribution to Meta Superintelligence Labs. See the upstream model card and LICENSE for terms, limitations, and responsible-use guidance.