xero0000/Gemma-4-26B-A4B-it-vram14-iq4xs-imat
Gemma-4-26B-A4B-it — vram14 (IQ4_XS + imatrix)
~13.3 GB all-VRAM mixed GGUF of google/gemma-4-26B-A4B-it (26B MoE, ~4B active), tuned so the full weight set fits dual mid-range GPUs and still beats the Q8 source on a coding/agent holdout.
TL;DR: 4.52 bpw custom mix (experts IQ4XS, attention Q5K, dense FFN Q6_K) + imatrix. On a 3060 Ti 8 GB + 3080 10 GB desktop: ~3060 t/s prefill / ~98 t/s decode at 128K all-VRAM — about 13.5× prefill and 5.3× decode vs the same model’s Q8 production recipe (CPU-MoE offload). Perplexity is quality-neutral vs Q8 on the target corpus.
File
Source was the official instruct Q8_0 GGUF (~25 GiB), requantized with an importance matrix (--allow-requantize).
Architecture notes
Gemma-4 26B-A4B GGUF layout (relevant for the recipe):
- 30 layers, 128 experts / 8 active
- Fused expert mat
ffn_gate_up_exps+ffn_down_expsdominate size (~90% of Q8 weights) - Tied embeddings (no separate
output.weight) ffn_down_expshas 704 columns → not 256-divisible; stock IQ4XS fails there, so that tensor uses **IQ4NL** instead
This arch currently needs a Gemma-4–capable llama.cpp build (TurboQuant / gemma4 fork or equivalent). Plain older mainline binaries that predate gemma4 will not load it.
Recipe
Built with llama-quantize (TurboQuant gemma4 tree), imatrix-guided:
Base ftype q4_K so custom --tensor-type rules apply. Approximate command:
llama-quantize --allow-requantize --imatrix gemma4-it.imatrix.gguf \
--token-embedding-type q6_K \
--tensor-type 'ffn_gate_up_exps=iq4_xs' \
--tensor-type 'ffn_down_exps=iq4_nl' \
--tensor-type 'attn_(q|k|v|output)\.weight=q5_K' \
--tensor-type 'ffn_(up|gate|down)\.weight=q6_K' \
--tensor-type 'ffn_gate_inp=q8_0' \
gemma-4-26B-A4B-it-Q8_0.gguf \
gemma-4-26B-A4B-it-vram14-iq4xs-imat.gguf \
q4_KImportance matrix: computed on a local coding/agent calibration corpus (~90 chunks, same domain family as the holdout). Imatrix here acts as light domain adaptation as well as rounding guidance.
Quality (perplexity)
Holdout corpus, n_ctx=512, 24 chunks (same methodology as the local vram13 series). Gemma PPL scale is not comparable to Qwen-family numbers (different tokenizer / corpus fit) — only within-family deltas matter.
Verdict on this holdout: 4-bit is quality-neutral to slightly better than Q8, because the imatrix + domain match more than offsets quantization noise. A stock Q4KM scored a bit lower PPL but was larger (15.6 GiB), needed CPU-MoE offload on this rig, and lost badly on runtime — it was deleted after the bench.
Runtime
Rig: RTX 3060 Ti 8 GB + RTX 3080 10 GB (18 GB total), Ryzen 5950X, DDR4. Engine: Gemma-4–capable llama.cpp (TurboQuant fork), q4_0 KV, flash-attn on. Perf prompt ≈ 6K-token prefill + 160 decode.
vram14 vs Q8 production: ~13.5× prefill, ~5.3× decode.
VRAM ceilings on 18 GB dual-GPU
At razor-edge VRAM the dual-GPU pipeline-parallel compute buffer reserve (~1.3–1.9 GiB on CUDA0) can log a transient allocation failure; the server then retries without pipeline parallelism and recovers at the speeds above. That log line is expected and harmless on this class of rig when fully loaded.
How to run
Thinking-capable instruct model — leave thinking enabled for agent/tool loops (official IT is trained to plan). Recommended sampling (Gemma defaults): temp 1.0, top_k 64, top_p 0.95. Clients that force temp ≈ 0.2 make tool loops pathologically deterministic.
Production (128K, all-VRAM)
./llama-server \
-m gemma-4-26B-A4B-it-vram14-iq4xs-imat.gguf \
--jinja --cache-type-k q4_0 --cache-type-v q4_0 --flash-attn on \
--ctx-size 131072 --parallel 1 --n-gpu-layers 99 \
--split-mode layer --tensor-split 44,56 \
--batch-size 2048 --ubatch-size 512 \
--temp 1.0 --top-k 64 --top-p 0.95 \
--no-mmap --threads 8 --no-warmup \
--port 8000Max context on 18 GB dual mid-range (192K, tight)
Same as above with --ctx-size 196608. Expect ~750 MiB free after settle — fine for interactive use, risky for long multi-tool sessions with a busy desktop compositor.
Single GPU
- ≥16 GB with headroom: drop
--tensor-split/ use a single-card split. - ≤12 GB: you will need expert offload (
--n-cpu-moe/ equivalent) and will lose most of the all-VRAM speedup; prefer a smaller quant or more VRAM.
Adjust --tensor-split for your card sizes (44,56 targets 8+10 GB).
Intended use & limitations
- Target: local chat / coding / agent workloads on ~16–18 GB total VRAM where you want Gemma-4 IT quality without Q8’s CPU-MoE tax.
- Multimodal (image) support depends on the runtime and GGUF export, not just weights — this release is validated as a text server quant.
- 4-bit experts are the quality floor vs full Q8; on this holdout the gap was closed by imatrix domain match, but other domains may differ.
- Inherits capabilities, refusal behavior, and biases of google/gemma-4-26B-A4B-it. No fine-tune — pure quantization.
Provenance
License: Apache-2.0 (same family as the base — Gemma 4 license). Quantization does not change the model license.
