Abiray/Gemma-4-12B-it-AEON-Abliterated-K4-GGUF
191.1k
Gemma-4-12B-it AEON Abliterated — K=4 GGUF Quants
This repository contains official GGUF quantizations of `AEON-7/Gemma-4-12B-it-AEON-Abliterated-K4-BF16`.
The base model is an abliteration of google/gemma-4-12B-it using a custom K=4 multi-direction norm-preserving biprojection that extends standard biprojection recipes with a K-dim orthonormal basis from the top-K SNR layers. This workflow preserves generative quality, slashes wikitext PPL drift compared to K=1 methods, and completely eliminates standard refusal patterns.
Provided Quantization Tiers
Deployment & Usage
1. Python (llama-cpp-python)
To run inference with full GPU acceleration, compile with the CUDA backend and load all layers into VRAM:
llama-cli \
--hf-repo Abhiray/Gemma-4-12B-it-AEON-Abliterated-K4-GGUF \
--hf-file Gemma-4-12B-it-AEON-Abliterated-Q4_K_M.gguf \
-ngl -1 \
-c 4096 \
-p "<start_of_turn>user\nYour prompt here<end_of_turn>\n<start_of_turn>model\n"