CoolFace
Modelpublic

mlasli/Muse-Glimmer-30B-Heretic-Abliterated-Q6_K-GGUF

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
1likes238downloads
Model Card

Muse Glimmer 30B - Heretic Abliterated (Q6_K GGUF)

v2 Release - Heretic-abliterated Muse Glimmer 30B in Q6_K GGUF format (~22 GB, very good quality).

Results

VersionRefusalsComplianceKL DivergenceTrials
v2 (current)6.5%93.5%0.076500
v129%71%0.02750

The v2 release achieves an 88% refusal reduction over v1.

Methodology

This model was abliterated using [Heretic](https://github.com/d3nd3/heretic) with 500 Optuna trials. See the BF16 model card for full methodology details.

Pipeline

  1. 1.Refusal directions computed from mlabonne/harmful_behaviors and mlabonne/harmless_alpaca
  2. 2.500 Optuna trials optimizing refusal vs. KL divergence
  3. 3.Best trial (Trial 445, 6.5% refusals, KL=0.076) applied via LoRA adapters
  4. 4.LoRA weights merged, then converted to GGUF with llama.cpp

GGUF Details

  • —Format: Q6_K
  • —File size: ~22 GB, very good quality
  • —Converted with: llama.cpp convert_hf_to_gguf.py
  • —Quantized with: llama.cpp llama-quantize

Usage

llama.cpp

bash
./llama-cli -m Muse-Glimmer-30B-Heretic-Abliterated-Q6_K.gguf -p "Your prompt here"

Ollama

Create a Modelfile:

dockerfile
FROM ./Muse-Glimmer-30B-Heretic-Abliterated-Q6_K.gguf

Then:

bash
ollama create muse-glimmer-30b-heretic-q6_k
ollama run muse-glimmer-30b-heretic-q6_k

Hardware Requirements

  • —RAM: ~22 GB, very good quality
  • —VRAM offloading: 12-24 GB recommended

Vision (Multimodal)

This model accepts image input when paired with a vision projector (mmproj). Abliteration only modified the language backbone — the vision encoder is untouched — so the standard Meta projector works directly with this repo.

This repository bundles mmproj-Muse-Glimmer-30B-Q4_K_M.gguf (~1.4 GB), Meta's official vision encoder

  • —projector for Muse Glimmer 30B.

Usage (llama.cpp)

bash
huggingface-cli download mlasli/Muse-Glimmer-30B-Heretic-Abliterated-Q6_K-GGUF \
  --include "Muse-Glimmer-30B-Heretic-Abliterated-Q6_K.gguf" \
  --include "mmproj-Muse-Glimmer-30B-Q4_K_M.gguf" \
  --local-dir ./models

./build/bin/llama-mtmd-cli \
  -m ./models/Muse-Glimmer-30B-Heretic-Abliterated-Q6_K.gguf \
  --mmproj ./models/mmproj-Muse-Glimmer-30B-Q4_K_M.gguf \
  --image photo.png \
  -p "Describe this image."
Ollama note: Ollama does not currently support separate mmproj files for this architecture. For image input, use llama.cpp (llama-mtmd-cli or llama-server --mmproj).

License

Apache 2.0 (same as base model)

Changelog

v1.1.0 — vision (multimodal) support (2026-08-16)

  • —Added mmproj-Muse-Glimmer-30B-Q4_K_M.gguf (~1.4 GB), Meta's official vision encoder + projector, enabling image input via llama.cpp.
  • —The vision tower is untouched by abliteration, so this projector matches the base model (meta-models/Muse-Glimmer-30B).
  • —v1.0.0 was the initial (unversioned) text-only upload.