XAEA12/Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliterated-GGUF
<div align="center">
<img src="./banner.svg" alt="Gemma 4 12B Coder — fable5 · composer2.5 · abliterated — GGUF" width="100%"/>
<h1>Huihui‑gemma‑4‑12B‑coder‑fable5‑composer2.5‑v1‑abliterated · GGUF</h1>
<p><b>GGUF quantization for fast local inference</b> of an abliterated (uncensored) coding model built on Gemma 4 12B.</p>
<p> <img alt="format" src="https://img.shields.io/badge/format-GGUF-22d3ee?style=for-the-badge"/> <img alt="quant" src="https://img.shields.io/badge/quant-Q4_K_M-a78bfa?style=for-the-badge"/> <img alt="params" src="https://img.shields.io/badge/params-12B-34d399?style=for-the-badge"/> <img alt="license" src="https://img.shields.io/badge/license-Gemma-f59e0b?style=for-the-badge"/> </p>
<p> <a href="#-quick-start-ollama">🚀 Quick start</a> · <a href="#-files">📦 Files</a> · <a href="#-conversion-details">🔧 Conversion</a> · <a href="#%EF%B8%8F-uncensored-model">⚠️ Notice</a> · <a href="#-license--credits">📜 Credits</a> </p>
</div>
✨ Overview
This repo provides a GGUF build of `huihui-ai/Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliterated` so you can run it locally with [Ollama](https://ollama.com) and [llama.cpp](https://github.com/ggml-org/llama.cpp).
The upstream model is a coding-focused fine-tune of Gemma 4 12B, trained on verifiable Python data with reasoning traces, then abliterated to remove refusals. Only the text tower is converted here (the multimodal vision/audio projector is omitted) — exactly what you want for code generation.
💡 The model emits a short reasoning pass (Thinking…) before the final answer, then the code.📦 Files
🚀 Quick start (Ollama)
Requires Ollama ≥ 0.30 (built-in gemma4 renderer/parser).
1. Download the GGUF and create a Modelfile:
FROM ./gemma4-coder-abliterated-Q4_K_M.gguf
TEMPLATE {{ .Prompt }}
RENDERER gemma4
PARSER gemma4
PARAMETER temperature 1
PARAMETER top_k 64
PARAMETER top_p 0.952. Build and run:
ollama create huihui-gemma4-coder-abliterated -f Modelfile
ollama run huihui-gemma4-coder-abliterated "Write a Python function for binary search."🦙 Quick start (llama.cpp)
llama-cli -m gemma4-coder-abliterated-Q4_K_M.gguf \
-p "Write a Python quicksort." -ngl 99 -c 8192🔧 Conversion details
- Converted from the original fp16
safetensorswith llama.cppconvert_hf_to_gguf.py(build b9775), then quantized withllama-quantize→ Q4_K_M. - Architecture:
Gemma4UnifiedForConditionalGeneration— text tower only. - Tokenizer conversion requires transformers ≥ 5.10 (older 4.x breaks on the Gemma 4 tokenizer config).
- ⚠️ The upstream model uses a
proportionalRoPE type on full-attention layers that currentllama.cppdoes not yet implement; it falls back to standard RoPE. Fine for typical use, but very long contexts (>32k) may differ from the originaltransformersbehavior.
⚠️ Uncensored model
This is an abliterated model — its safety filtering has been significantly reduced and it may generate sensitive, controversial, or otherwise inappropriate content. Use responsibly, review outputs, and you are solely responsible for your use. This is provided as-is for research and local experimentation.
📜 License & credits
Derivative of a Gemma model, distributed under the [Gemma Terms of Use](https://ai.google.dev/gemma/terms) (including the Prohibited Use Policy). All credit for the model itself goes to:
This repository contributes only the GGUF quantization for local inference.
<div align="center"><sub>Made for local inference with ❤️ · Ollama & llama.cpp</sub></div>
