deucebucket/Gemma-4-26B-A4B-it-Cerebellum-v6-GGUF
<p align="center"> <img src="cerebellum_banner.png" alt="Cerebellum" width="640"> </p>
Gemma 4 26B-A4B-it Cerebellum GGUF
Sensitivity-guided mixed-precision GGUF of google/gemma-4-26B-A4B-it: a Q3KM base with the Cerebellum v6 tensor allocation. The shipped file carries the v6 weights plus Google's updated Gemma 4 chat-template metadata (the 2026-05-18 template state) with zero tensor changes versus v6. Newer versions appear in filenames, not the repo name.
Files
Evaluation
Measured directly on the GGUF with llama.cpp llama-server on an RTX 3090, temperature 0, project benchmark harness. v6.1 is metadata-only over v6, so these describe the same weights. The comparison column is our own same-size uniform Q3KM build measured on the same harness. Summary JSONs are in benchmark_results/.
HumanEval for Gemma 4 must use the chat-completions harness (scripts/benchmark_evalplus_chat.py, enable_thinking: false, thinking_budget_tokens: 0, BENCH_WORKERS=1). The retained v6 HumanEval artifacts were raw-completions and are marked for re-audit, so no v6 HumanEval number is published here.
Usage
Gemma 4 requires --jinja. For non-thinking output, pass request-level chat_template_kwargs: {"enable_thinking": false} and thinking_budget_tokens: 0; do not set a fixed server --reasoning-budget (it can burn output into hidden reasoning until the length cap, which looks like a repetition loop).
llama-server \
--model gemma-4-26B-A4B-it-cerebellum-v6.1-templatefix-Q3_K_M.gguf \
--mmproj gemma-4-26b-a4b-it.mmproj.gguf \
-ngl 99 --ctx-size 65536 --parallel 1 --flash-attn on \
--cache-type-k q8_0 --cache-type-v q8_0 --jinja --reasoning autoMeasured on one RTX 3090 (24 GB), KV q8_0: ~123 tok/s decode, 15.1 GB peak VRAM (4-slot serving), context to 131,072. This rig's measurements; no quality claims beyond them.
Provenance
- Base: google/gemma-4-26B-A4B-it — Google Gemma Team
- Base quant lineage: Q3KM with the bartowski imatrix (
bartowski/google_gemma-4-26B-A4B-it-GGUF) - Recipe: Cerebellum v6 tensor allocation; v6.1 is a chat-template metadata refresh (Google 2026-05-18 template), zero tensor changes
Credits
- Base model: Google Gemma Team,
google/gemma-4-26B-A4B-it - Imatrix: bartowski,
bartowski/google_gemma-4-26B-A4B-it-GGUF - GGUF runtime: llama.cpp
- Quantization method: Cerebellum — deucebucket
