roman220220/gemma-4-26B-A4B-it-GGUF-jang-imatrix
<p align="center"> <img src="llmtray-banner.png" alt="LLMTray" width="100%"> </p>
Gemma 4 26B-A4B (MoE) — GGUF, imatrix + JANG-mixed, text + vision
### ▶ GGUF build for ollama / llama.cpp This is the GGUF build (ollama / llama.cpp). If you use LLMTray — IPSupport's local AI app for Apple Silicon, which runs MLX — use the MLX sibling roman220220/gemma-4-26B-A4B-it-gptq-mlx-jang instead.  
A GGUF quantization of google/gemma-4-26B-A4B-it (the 128-expert / top-8 MoE variant, ~4B active params) for llama.cpp / Arc / CPU / any GGUF runtime — with two deliberate improvements over a stock Q3_K_M:
- imatrix computed on a broad real corpus (bartowski's
calibration_datav3— code + prose + facts + multilingual). This is the GGUF-world equivalent of calibrating on real data, and lets a smaller quant match the quality of a larger stock one. - JANG-mixed precision (this account's "spend bits where they matter" recipe): attention kept HIGH (
Q5_K), the 128 routed experts — which are ~90% of the weights — pushed LOW (IQ3_S). Same principle as the MLX JANG releases, expressed viallama-quantize --tensor-typeoverrides.
Result: 12.5GB text model (smaller than a stock ~13GB Q3KM) plus a quantized vision projector — the stock text-only q3km has no vision.
Files
Verified working (text + vision)
Text (chat):
Here are three facts about the Roman Empire: 1. It reached its greatest territorial extent under Emperor Trajan (117 AD)... Mare Nostrum... 2. over 400,000 km of roads...
Vision (real COCO photo of two cats):
The picture shows two cats sleeping on a couch.
Both run through this actual quantized checkpoint, ~130 tok/s generation on an A100.
Usage (Ollama) — use the included Modelfile
Plain ollama run hf.co/roman220220/gemma-4-26B-A4B-it-GGUF-jang-imatrix loads and runs, but thinking/tool tokens leak into the chat as raw text (<|channel>thought … <channel|>). A Hugging Face repo can only give Ollama a Go template / params / system prompt — it cannot set Ollama's native Gemma 4 RENDERER/PARSER, which is what the official gemma4 model uses to hide the thinking block and parse Gemma 4's native tool-call syntax. The included Modelfile sets both and pulls the weights straight from this repo:
curl -L -o Modelfile https://huggingface.co/roman220220/gemma-4-26B-A4B-it-GGUF-jang-imatrix/resolve/main/Modelfile
curl -L -o gemma4-26b-assistant-q8_0.gguf https://huggingface.co/roman220220/gemma-4-26B-A4B-it-GGUF-jang-imatrix/resolve/main/gemma4-26b-assistant-q8_0.gguf
ollama create gemma4-26b-jang -f Modelfile
ollama run gemma4-26b-jangSpeculative decoding (MTP drafter). gemma4-26b-assistant-q8_0.gguf (462MB) is Google's Multi-Token-Prediction drafter for Gemma 4 26B-A4B (google/gemma-4-26B-A4B-it-assistant), byte-identical to the draft layer of the official ollama gemma4:26b (sha256 6326fb9f…). It's a 4-layer head that reads the main model's hidden state and KV cache and guesses a few tokens ahead; the main model verifies them in one pass, so output is identical and decode is faster. Ollama's DRAFT only takes a local file, hence the second download. The drafter was trained against the bf16 model, so acceptance on this quantized one may be a bit lower than Google's "up to 3x" — measure tok/s with and without it. If memory is tight, delete the two draft lines from the Modelfile.
Requires a recent Ollama (Gemma 4 support; the official model declares requires: 0.30.0).
Usage (llama.cpp) — --jinja is REQUIRED
Gemma 4's chat template uses control tokens (<|turn>, <|channel>, <|think|>); without --jinja, llama.cpp applies a simplified template and chat/thinking mode breaks (the model emits raw <thought and stops). Always pass --jinja.
Text:
llama-cli -m gemma4-26b-a4b-jang-iq3s.gguf -ngl 99 --jinja \
-p "List three facts about the Roman Empire."Vision:
llama-mtmd-cli -m gemma4-26b-a4b-jang-iq3s.gguf \
--mmproj mmproj-gemma4-26b-q8.gguf -ngl 99 --jinja \
--image photo.jpg -p "What is in this picture?"Honest notes
- This is imatrix + JANG bit-allocation, NOT this account's GPTQ. GPTQ's Hessian weight-correction is calibrated against a specific quant grid; llama.cpp's k-quant super-block grid differs, so the correction doesn't transfer and gets re-quantized away. imatrix + JANG are the parts that DO carry over to GGUF. (The GPTQ version of this model lives at roman220220/gemma-4-26B-A4B-it-gptq-mlx-jang, for MLX / Apple Silicon.)
- Perplexity as a metric is currently broken for Gemma 4 MoE in llama.cpp — both this quant and stock reference GGUFs report absurd PPL (~25k–33k on wikitext) while generating perfectly. Quality here was verified by generation, not PPL.
- Requires a recent llama.cpp built from source (Gemma 4 MoE + vision support landed only in the
conversion/gemma.pyrefactor; prebuilt release binaries don't have it yet).
MLX / Apple Silicon: an 8-bit MLX build of the same drafter is at roman220220/gemma-4-26B-A4B-it-assistant-mlx-8bit.
Method / code
- Pipeline: rromenskyi/quant-ternary/gemma4-quant (
gemma4_26b_gguf_pipeline.sh,gemma4_26b_fix_tokenizer.py)
The IPSupport local-AI stack
Local-first AI tools for macOS by IPSupport — nothing leaves your Mac.
- [LLMTray](https://www.ipsupport.us/llmtray/) — your local AI workstation for macOS: chat with local LLMs, generate and edit images, make music, run agents, and serve an OpenAI-compatible API. Downloads models from Hugging Face in-app, with per-model profiles.
- [IPSupport Code](https://ipsupport-llc.github.io/ipsupport-code/) — your AI coding agent for real repositories: analyze, fix, test, report.
This is the GGUF build (ollama / llama.cpp). For the MLX build LLMTray runs, see roman220220/gemma-4-26B-A4B-it-gptq-mlx-jang.
License
Licensed under the Apache License 2.0, the same license as the base model — see `LICENSE`.
Modified from google/gemma-4-26B-A4B-it: converted to GGUF and quantized with an importance matrix and per-tensor types (attention Q5_K, dense MLP Q3_K_M, routed experts IQ3_S); the vision tower and projector were converted to a separate Q8_0 GGUF (some tensors F16). The weights and configuration files in this repo are therefore modified versions of the original, not the original files.
gemma4-26b-assistant-q8_0.gguf is Google's MTP drafter google/gemma-4-26B-A4B-it-assistant (also Apache 2.0) converted to GGUF Q8_0.
Gemma 4 is released by Google under Apache 2.0 (Gemma 4 license terms).
