Vishva007/gemma-4-E4B-it-W4A16-AutoRound-AWQ
Gemma 4 E4B Instruct — W4A16 Quantized AWQ
This repository hosts W4A16 INT4-quantized versions of `google/gemma-4-E4B-it`, a multimodal mixture-of-experts model supporting text, vision, and audio inputs. Two quantized variants are available:
Note on MTP / Speculative Decoding: If you want to use the official speculative decoding assistant model (google/gemma-4-E4B-it-assistant) for MTP support, it is recommended to use thevllm/vllm-openai:gemma4-0505-cu129Docker image, which includes newer Gemma 4 support and decoding patches. Due to INT4 quantization, the assistant acceptance rate may be lower compared to the original unquantizedgoogle/gemma-4-E4B-itmodel.
Note: Only the language model (LM) layers are quantized to INT4. The vision tower, audio tower, and multimodal projectors are kept at full precision (BF16) to preserve multimodal quality.
Quantization Details
Model Architecture
Gemma 4 E4B is a multimodal MoE model (Gemma4ForConditionalGeneration) with:
- Text backbone: 42-layer
Gemma4TextModelwith 2560 hidden dim, mixed local/global attention - Vision tower: 16-layer
Gemma4VisionModel(768-dim, unquantized) - Audio tower: 12-layer
Gemma4AudioModelwith conformer-style layers (unquantized) - Vocabulary size: 262,144 tokens
Usage
vLLM Inference
The recommended way to serve this model is via the official `vllm/vllm-openai:gemma4` Docker image, which ships vLLM v0.19.1 with the latest Transformers patches required for Gemma 4.
Serve with Docker (recommended)
docker run --gpus all --rm -p 8000:8000 \
vllm/vllm-openai:gemma4 \
vllm serve Vishva007/gemma-4-E4B-it-W4A16-AutoRound-AWQ \
--served-model-name Gemma-4-E4B-it \
--quantization autoround \
--kv-cache-dtype auto \
--max-num-batched-tokens 16384 \
--enable-chunked-prefill \
--enable-prefix-caching \
--dtype bfloat16 \
--max-model-len 18432 \
--enable-auto-tool-choice \
--tool-call-parser gemma4 \
--port 8000 \
--default-chat-template-kwargs '{"enable_thinking": false}' \
--mm-processor-kwargs '{"max_soft_tokens": 560}'Direct vllm serve (vLLM ≥ 0.19.0)
vllm serve Vishva007/gemma-4-E4B-it-W4A16-AutoRound-AWQ \
--served-model-name Gemma-4-E4B-it \
--quantization autoround \
--kv-cache-dtype auto \
--max-num-batched-tokens 16384 \
--enable-chunked-prefill \
--enable-prefix-caching \
--dtype bfloat16 \
--max-model-len 18432 \
--enable-auto-tool-choice \
--tool-call-parser gemma4 \
--port 8000 \
--default-chat-template-kwargs '{"enable_thinking": false}' \
--mm-processor-kwargs '{"max_soft_tokens": 560}'max_soft_tokens — Image Token Budget
The max_soft_tokens parameter controls how many visual tokens are allocated per image. Higher values give richer image representations at the cost of context length and throughput.
Pass it via --mm-processor-kwargs '{"max_soft_tokens": <value>}'.
OpenAI-compatible API call
Once the server is running, query it like any OpenAI-compatible endpoint:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="Gemma-4-E4B-it",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Describe what you see."},
{"type": "image_url", "image_url": {"url": "https://..."}},
],
}
],
max_tokens=512,
)
print(response.choices.message.content)The quantized model was then exported in both AutoRound and AWQ formats and pushed to Hugging Face Hub.
Limitations & Notes
- RTN mode (
iters=0) is used instead of full AutoRound optimization due to Gemma 4's architecture constraints. - Some layers with shapes not divisible by 32 are skipped during quantization (minor precision impact).
- Multimodal (vision/audio) capabilities are fully preserved as those towers are not quantized.
🚀 Deploy on RunPod
One-click launch environments pre-configured with PyTorch, CUDA, and dependencies for fine-tuning or quantization.
🎁 Need GPU compute? Sign up via RunPod and get $5–$500 in free credits when you add your first $10.
PyTorch 2.14
PyTorch 2.13
PyTorch 2.12
Acknowledgements
Special thanks to OLAF-OSS and the contributors of the gemma4-vllm project for their excellent work in enabling Gemma 4 support with vLLM. Their QUANTIZE.md guide was extremely helpful in understanding the correct quantization approach for gemma-4-E2B-it.
At the time of writing, the original repository appears to be unavailable or removed. To ensure reproducibility, I have documented the full quantization workflow used for this model in my own repository.
The full quantization process used to produce these models is documented here: 📓 auto_round_Gemma4-E4B.ipynb
License
This quantized model is derived from `google/gemma-4-E4B-it` and is subject to the Gemma Terms of Use.
Citation
If you use this quantized model, please also cite the original Gemma 4 work:
@misc{gemma4_2026,
title = {Gemma 4},
author = {Google DeepMind},
year = {2026},
url = {https://huggingface.co/google/gemma-4-E4B-it}
}Quantized by [Vishva007](https://huggingface.co/Vishva007)
