culturerevolt/gemma-4-12b-heretic-abliterated-GGUF
Gemma-4-12B-Heretic-Abliterated-iMatrix-GGUF (Complete Suite)
This repository hosts a complete, professional-grade suite of iMatrix (Importance Matrix) GGUF quantizations of culturerevolt/gemma-4-12b-heretic-abliterated.
The base model is an abliterated, fully decensored variant of Google's gemma-4-12b-it architecture, stripped of categorical refusal strings via norm-preserving directional ablation. This comprehensive GGUF suite is calibrated to preserve maximum linguistic entropy, narrative depth for creative fiction, and strict syntax parsing for agentic tools across multiple hardware profiles.
๐ Quick-Reference: Which Quant Should You Download?
All models (except the standard Q8_0 baseline) utilize a custom calculated importance matrix to prevent the loss of reasoning performance that standard static quants suffer from. Use this table to match the model files to your available hardware footprint:
๐ง Technical Methodology: The iMatrix Advantage
When compressing a model down to 3, 4, or 5 bits, a basic, uncalibrated static quantization acts like a global color reduction on an image because it destroys critical structural details. Large Language Models contain sensitive outlier weights that dictate logic patterns, punctuation structure, and instruction adherence. Indiscriminately compressing these weights leads to specialization collapse, which causes a model to loop tokens, output formatting gibberish, or hallucinate heavily.
To protect this repo from degradation, these files were compiled using a Custom Hybrid Importance Matrix (`imatrix.dat`) strategy.
1. Calibration Profile
The calibration matrix was generated over a high-entropy, multi-domain text dataset specifically balanced for the local LLM stack:
- Core Logic & Reasoning: Grounded using community standard English linguistic corpuses (
text_en_large) to lock down general intelligence, deductive consistency, and vocabulary breadth. - Agentic Punctuation: Interwoven with structured tool-calling strings (
tools_large) to safeguard brackets, colons, nested variables, and JSON structures necessary for automated file parsing pipelines like VaultForge. - Narrative Fidelity: Infused with creative prose samples to explicitly train the importance matrix to prioritize complex character interiority and vivid world-building pacing.
2. Execution Parameters
To avoid calculation errors or PCIe data-transfer distortions, the matrix was evaluated over a tight, isolated memory profile:
- Total Cycles: 200 high-entropy chunks (representing roughly 100,000 tokens of diverse data)
- Chunk Boundary: 512 tokens
- Synchronization Sizing:
-b 512 -ub 512(Forcing physical batch alignment to protect mathematical scaling)
The resulting map acts as an explicit instruction booklet during quantization. It forces the llama-quantize tool to fiercely protect sensitive reasoning neurons while aggressively compressing the easy background language weights.
โ๏ธ Recommended Inference Settings
The Gemma-4 family is a highly capable but precise architecture. To avoid text stuttering, formatting drops, or interface crashes in backends like LM Studio, AnythingLLM, or llama-server, implement the following configurations:
1. Multi-Modal Vision Execution
Gemma-4 features a unified, encoder-free architecture, projecting raw image patches straight into the embedding layers. To handle vision tasks inside llama.cpp based frontends, you must load the companion multimodal projector file alongside your chosen text quant.
The verified companion file gemma-4-12b-it-mmproj-f16.gguf is hosted right here in the repository root. Ensure you map this file inside your backend Vision Adapter settings slot to seamlessly initialize image ingestion.
๐ Acknowledgements
- Google DeepMind for pioneering the unified Gemma-4 architecture.
- Philipp Emanuel Weidmann for developing the underlying Heretic abliteration framework.
- Massive thanks to the open-source local AI community for continuously pushing the boundaries of what is possible on local consumer hardware.
โ ๏ธ Disclaimer & Boundary Limits
This model is completely unaligned. It will output text without filtering, judgment, or warning labels. By downloading this model, you accept full responsibility for the prompts fed to it and the text generated by it. Use responsibly within local sandbox development setups.
๐๏ธ Streamlined Jinja Chat Template
If you encounter interface parsing issues with heavy multi-turn configurations or want to maximize token efficiency during rapid back-and-forth chat sessions, use this clean, hyper-efficient template:
{%- for message in messages %}
<|turn|>{{ message['role'] }}
{{ message['content'] }}<|turn|>
{%- endfor %}
{%- if add_generation_prompt %}
<|turn|>assistant<|channel>thought <channel|>
{%- endif %}