CoolFace
Modelpublic

Abiray/Gemma-4-12B-it-AEON-Abliterated-K4-GGUF

sourceHugging Facegemmaupdated 4mo agoView on Hugging Face
19likes1.1kdownloads
Model Card

Gemma-4-12B-it AEON Abliterated — K=4 GGUF Quants

This repository contains official GGUF quantizations of `AEON-7/Gemma-4-12B-it-AEON-Abliterated-K4-BF16`.

The base model is an abliteration of google/gemma-4-12B-it using a custom K=4 multi-direction norm-preserving biprojection that extends standard biprojection recipes with a K-dim orthonormal basis from the top-K SNR layers. This workflow preserves generative quality, slashes wikitext PPL drift compared to K=1 methods, and completely eliminates standard refusal patterns.

Provided Quantization Tiers

File NameQuant TypeSizeVRAM / Hardware Recommendation
Gemma-4-12B-it-AEON-Abliterated-Q3_K_M.gguf3-bit6.09 GBLow-VRAM setups / CPU+GPU split
Gemma-4-12B-it-AEON-Abliterated-Q4_K_M.gguf4-bit7.38 GBRecommended balance. Fits entirely on single T4/RTX 3060/4060
Gemma-4-12B-it-AEON-Abliterated-Q5_K_M.gguf5-bit8.55 GBLow quality degradation, tight fit on 8GB VRAM
Gemma-4-12B-it-AEON-Abliterated-Q6_K.gguf6-bit9.79 GBHigh fidelity, excellent for 12GB+ VRAM
Gemma-4-12B-it-AEON-Abliterated-Q8_0.gguf8-bit12.70 GBNear-identical to BF16 precision. Fits comfortably on 16GB VRAM

Deployment & Usage

1. Python (llama-cpp-python)

To run inference with full GPU acceleration, compile with the CUDA backend and load all layers into VRAM:

bash
llama-cli \
  --hf-repo Abhiray/Gemma-4-12B-it-AEON-Abliterated-K4-GGUF \
  --hf-file Gemma-4-12B-it-AEON-Abliterated-Q4_K_M.gguf \
  -ngl -1 \
  -c 4096 \
  -p "<start_of_turn>user\nYour prompt here<end_of_turn>\n<start_of_turn>model\n"