rajasingh012/gemma-4-12b-it-quark-w8a8-int8
Gemma-4-12B-it-Quark-W8A8-INT8
W8A8 INT8 quantized version of google/gemma-4-12b-it using AMD Quark, produced as part of the RetailConcierge AMD AI DevMaster 2026 submission.
Model Details
Quantization Scheme
How to Use
With vLLM (Recommended)
# Start the server (single AMD MI300X is enough)
vllm serve rajasingh012/gemma-4-12b-it-quark-w8a8-int8 \
--tensor-parallel-size 1 \
--max-model-len 32768 \
--gpu-memory-utilization 0.9 \
--quantization quark \
--trust-remote-code \
--enable-prefix-caching \
--enable-chunked-prefill
# Chat completion
curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{
"model": "rajasingh012/gemma-4-12b-it-quark-w8a8-int8",
"messages": [{"role": "user", "content": "Recommend a pair of wireless earbuds under 5000 rupees."}],
"max_tokens": 512,
"temperature": 0.7
}'Requires vLLM >= 0.26 (the --quantization quark loader). Tested on vLLM 0.26.0+rocm723 (AMD ROCm wheel) with --tool-call-parser gemma4.
Hardware Requirements
- Minimum: 1× GPU with ≥48 GB VRAM (e.g., AMD MI300X / MI350X, NVIDIA A100-80G).
- Quantized weights measure ~12.5 GiB on GPU, leaving ample KV cache headroom.
Quantization Details
This model was quantized using AMD Quark's per-token per-channel INT8 scheme (W8A8):
- Weight quantization: INT8 per-channel (one scale per output channel), symmetric, static.
- Activation quantization: INT8 per-token (one scale per token), symmetric, dynamic (computed at inference time, so no calibration data needed).
- Excluded layers:
model.vision_embedder.*,model.embed_vision.embedding_projection,model.embed_audio.embedding_projection(vision/audio modalities preserved in BF16).lm_headshares storage withembed_tokens(tie_word_embeddings: True). - Export: real INT8 weights with BF16 scales (no fake-quant, no zero-point), Quark 0.12
export_safetensors. - Post-export fixup (required): Quark 0.12 exports some vision keys under names vLLM's
gemma4_unifiedloader does not map (embed_vision.multimodal_embedder.*,embed_vision.patch_*). The fixup renames them to the expected layout and copieschat_template.jinjainto the output (vLLM 400s chat requests without it). The fixup is automated in the RetailConcierge repo (scripts/_quark_fix_vllm_keys.py) and is required for any Gemma 4 Unified checkpoint.
Reproduce Quantization
Full reproducible recipe in the RetailConcierge repo, scripts/ folder:
# On the AMD GPU droplet (see scripts/README.md for full lifecycle)
bash /root/quantize_int8.sh --model google/gemma-4-12b-itThe script handles: preflight (Quark version, transformers >= 5.10.1 for Gemma4UnifiedForConditionalGeneration, BF16 source location), the Quark W8A8 quantize, the post-export key fixup, and output verification. Recipe provenance: nameistoken/Gemma-4-31B-it-Quark-W8A8-INT8 (−0.08pp GSM8K on the 31B dense class); the 12B Unified run is a fresh quantization of the same scheme.
Accuracy
Caveat: the −0.08pp GSM8K figure published on the 31B dense baseline was measured on the older Gemma4ForConditionalGeneration class. The 12B Unified class uses a different multimodal embedding layout; accuracy for this checkpoint is validated via the RetailConcierge tool-call accuracy gate (extractbrief → searchcatalog → finalize_recommendations pass rate vs BF16), not yet via GSM8K. Treat as a fresh quantization until independently benchmarked.
Citation
If you use this model, please cite the original Gemma 4 release:
@misc{google2026gemma4,
title = {Gemma 4},
author = {Google DeepMind},
year = {2026},
url = {https://huggingface.co/google/gemma-4-12b-it}
}License
This model is released under the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0), following the Gemma 4 license under which the upstream google/gemma-4-12b-it weights are distributed by Google DeepMind.
This is a quantized derivative of google/gemma-4-12b-it. Per Apache 2.0 §4:
- Modified files (the INT8-quantized
model.safetensorsand the appendedquantization_configblock inconfig.json) carry this notice as part of the model card. - Original copyright and attribution notices from the base model are preserved (see
NOTICE). - A copy of the Apache 2.0 license text is included as
LICENSE.
Original weights © Google DeepMind. Quantization performed by the model author; no warranty of any kind is provided (see LICENSE §7–8).
