roman220220/gemma-4-26B-A4B-it-gptq-mlx-jang
<p align="center"> <img src="llmtray-banner.png" alt="LLMTray" width="100%"> </p>
Gemma 4 26B-A4B (MoE) — GPTQ, JANG-mixed precision, text + vision, MLX
### ▶ Run it locally in LLMTray A free, native macOS app for local AI on Apple Silicon — chat, images, music, agents and an OpenAI-compatible API. Point it at this model; nothing leaves your Mac.  
A JANG-style mixed-precision MLX quantization of google/gemma-4-26B-A4B-it — the 26.5B-parameter, 128-expert / top-8 MoE variant of Gemma 4 (≈4B active parameters per token, hence "A4B") — using GPTQ Hessian-based error correction across every quantizable Linear in both the text decoder and the vision tower. Attention projections at 8-bit, feed-forward (dense MLP and all 128 routed experts) at 4-bit, embeddings at 8-bit RTN. Nothing is left in bf16.
- ~15GB (down from 51.6GB bf16 — 3.4× smaller).
- This variant has no audio tower (
audio_configis null on the real checkpoint) — text + vision only, unlike the E4B sibling.
Recipe: JANG-mixed, by role
Calibrated with --group-size 64 on real data: diverse text prompts for the language model, real COCO photographs for the vision tower.
Architecture notes (verified directly against the real checkpoint)
- Hybrid dense+MoE layers: every one of the 30 decoder layers computes BOTH a dense MLP path and a routed-experts path, combined additively (
hidden_states_1 + hidden_states_2) — the router operates on the pre-dense-MLP residual. Both paths are quantized here. - Per-expert GPTQ calibration: the
Gemma4TextExpertsmodule consumes raw stacked[128, out, in]nn.Parametertensors inside a Python loop, notnn.Linearsubmodules — so calibration hooks the whole experts module, replicates its realexpert_mask/token_idxgathering to bucket calibration rows per expert, and runs batched GPTQ across all 128 experts at once.down_proj's calibration input is recomputed asact_fn(gate)*upusing the real (uncorrected)gate_up_proj, exactly as the real forward pass feeds it. - Vision intermediate_size padding: the vision MLP's intermediate width (4304) isn't divisible by any group size
mx.quantizesupports (32/64/128 — 4304 = 2⁴·269). Rather than leave those tensors in bf16,gate/up/down_projare zero-padded to 4352 (next multiple of 64) — this is mathematically exact (GELU(0)·0 = 0, and the padded activations hit correspondingly-zero-weighteddown_projcolumns), declared inconfig.json'svision_config.intermediate_sizeso the model is built at the padded width with no inference-code changes.
Does it work? Real end-to-end validation
Text: correctly answers factual/reasoning prompts (Gemma 4's thinking-mode chain included).
Vision (real COCO photo of two cats, "What animal is in this picture?"):
There are two cats in this picture.
Both run through this actual GPTQ-quantized checkpoint (MoE routing + vision tower + fusion all exercised).
Usage
pip install git+https://github.com/ipsupport-llc/mlx-lm.git@mainfrom mlx_lm import load, generate
model, tok = load("roman220220/gemma-4-26B-A4B-it-gptq-mlx-jang")
messages = [{"role": "user", "content": "What is the capital of France?"}]
print(generate(model, tok, prompt=tok.apply_chat_template(messages, add_generation_prompt=True)))Image input needs a manual generation loop (mlx_lm.generate() has no image plumbing for this fork's vision models yet) — full example in the gemma4-quant pipeline repo.
Method / code
- Model code: ipsupport-llc/mlx-lm (
gemma4.py,gemma4_text.py,gemma4_vision.py) - Quantization pipeline: rromenskyi/quant-ternary/gemma4-quant (
gemma4_26b_gptq_calibrate.py,gemma4_26b_gptq_splice.py) - Same
gptq_nbit/gptq_nbit_batchedimplementation as this account's other quantization work (nemotron-extreme-quant,zimage-quant), unmodified.
The IPSupport local-AI stack
Local-first AI tools for macOS by IPSupport — nothing leaves your Mac.
- [LLMTray](https://www.ipsupport.us/llmtray/) — your local AI workstation for macOS: chat with local LLMs, generate and edit images, make music, run agents, and serve an OpenAI-compatible API. Downloads models from Hugging Face in-app, with per-model profiles.
- [IPSupport Code](https://ipsupport-llc.github.io/ipsupport-code/) — your AI coding agent for real repositories: analyze, fix, test, report.
Runs in LLMTray as a chat model, with vision.
License
Licensed under the Apache License 2.0, the same license as the base model — see `LICENSE`.
Modified from google/gemma-4-26B-A4B-it: the text decoder and vision tower were GPTQ-quantized (attention 8-bit; dense MLP and all 128 MoE experts 4-bit; group size 64), embeddings were quantized to 8-bit RTN, the router was left untouched, the vision MLP was zero-padded from 4304 to 4352 (with config.json updated to match), and the model was converted to MLX. The weights and configuration files in this repo are therefore modified versions of the original, not the original files.
Gemma 4 is released by Google under Apache 2.0 (Gemma 4 license terms).
