lemuralabs/Gemma-4-12B-uncensored-bf16
<p align="center"> <img src="logo.png" alt="Lemura Labs" width="110"/> </p>
Gemma-4-12B-uncensored-bf16
Full-precision (bf16) abliterated google/gemma-4-12B-it — the complete encoder-free unified multimodal model (text · image · audio · video) with refusals removed via the ablation toolkit. *This is the artifact that runs refusal-free vision + audio + video today* (in transformers), and the source for the MLX quants below. By Lemura Labs**.Abliterated model — read this
Refusal directions were surgically removed from the parent. It will answer many prompts the parent refuses. No new capabilities were added — only refusal behavior was reduced. Use responsibly and within applicable law.
Refusal removal — before / after
Measured with the ablation toolkit's evaluator on 100 harmful prompts (mlabonne/harmful_behaviors test[:100]), greedy decoding, refusal-marker classifier:
87 fewer refusals — an 87.9% reduction, at KL divergence 0.053 from the original (≪ 0.5, the damage threshold) → general capabilities preserved.
Specs
Inference & compatibility
Quick start — transformers (text)
pip install -U "transformers>=5.10" torch torchvision librosa acceleratefrom transformers import AutoProcessor, AutoModelForMultimodalLM
mid = "lemuralabs/Gemma-4-12B-uncensored-bf16"
processor = AutoProcessor.from_pretrained(mid)
model = AutoModelForMultimodalLM.from_pretrained(mid, dtype="auto", device_map="auto")
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain abliteration in two sentences."},
]
inputs = processor.apply_chat_template(messages, tokenize=True, return_dict=True,
return_tensors="pt", add_generation_prompt=True, enable_thinking=False).to(model.device)
n = inputs["input_ids"].shape[-1]
out = model.generate(**inputs, max_new_tokens=256)
print(processor.parse_response(processor.decode(out[0][n:], skip_special_tokens=False)))enable_thinking=Trueturns on reasoning mode;parse_responseseparates the thinking channel.
Vision & audio (image · audio · video)
Full multimodal runs here today — pass image/audio/video in the message content:
messages = [{"role": "user", "content": [
{"type": "image", "url": "https://.../photo.jpg"}, # image → key "url"
{"type": "audio", "audio": "https://.../clip.wav"}, # audio → key "audio" (≤30s)
{"type": "text", "text": "Describe what you see and hear."},
]}]
inputs = processor.apply_chat_template(messages, tokenize=True, return_dict=True,
return_tensors="pt", add_generation_prompt=True).to(model.device)
n = inputs["input_ids"].shape[-1]
out = model.generate(**inputs, max_new_tokens=512)
print(processor.parse_response(processor.decode(out[0][n:], skip_special_tokens=False)))Audio ≤ 30 s (native ASR + speech translation) · images variable-resolution · video ≤ 60 s (~1 fps).
Running on Mac
This bf16 repo runs in transformers on Apple Silicon (MPS) — full multimodal, as above. For lighter, faster MLX serving, use the MLX quants of this model (see the family table) with: **oMLX** (inference server + macOS menu-bar app, SSD KV cache), **vMLX**, **LM Studio** (MLX engine), **Ollama** 0.19+, or **mlx-vlm** directly. Those serve the MLX quants once their bundled mlx-lm/mlx-vlm adds gemma4_unified support (text today via mlx-vlm + a small shim).
Quant family
Lineage
google/gemma-4-12B (Google DeepMind — base pretrain)
↓ instruction tuning
google/gemma-4-12B-it (multimodal, encoder-free)
↓ ablation 1.3.0 — directional ablation, Optuna/TPE-optimized over 100 trials, best Pareto trial #55
this repo — abliterated bf16 (refusals 99→12 / 100, KL 0.053)
↓ mlx-vlm quantization
MLX quants (8-bit · MXFP4 · mixed) — see family tableCredits
License
Apache-2.0 (inherited from the base). Also subject to the Gemma 4 Terms of Use.
