CoolFace
Modelpublic

positron-ai/google_gemma-4-26B-A4B-it-ingest-best-gptq

sourceHugging Faceapache-2.0updated 26d agoView on Hugging Face
0likes530downloads
Model Card

Positron AI Quantized Build

This repository contains a Positron AI quantized build of google/gemma-4-26B-A4B-it for inference.

Recommended Use

Use this artifact when you need a GPTQ 4-bit build of google/gemma-4-26B-A4B-it built by Positron AI.

For general-purpose GPU inference, compare against the original model and other quantized formats before deployment.

Artifact Summary

FieldValue
Base modelgoogle/gemma-4-26B-A4B-it
Published artifactpositron-ai/google_gemma-4-26B-A4B-it-ingest-best-gptq
Quantization methodGPTQ
Quantization formatgptq
Source precisionn/a
Target runtimen/a
Hardware targetn/a
Release date2026-09-01
Licenseapache-2.0

Quantization Details

FieldValue
Weight precision4-bit
Activation precisionnot quantized
Bits4
Group size64
Symmetric quantizationtrue
Activation ordering / desc_actfalse
Damp percent0.05
Calibration datasetMixed-domain calibration set
Calibration samples128
Calibration sequence length4096
MoE experts per tokenn/a
Quantization toolchainGPTQModel 7.2.0, transformers 5.11.0, torch 2.9.1, CUDA 12.8

Evaluation

This card intentionally reports no performance or quality metrics (no KL-divergence, accuracy, or perplexity figures). Validation results are tracked internally by Positron AI.

Provenance

This artifact was produced by Positron AI from google/gemma-4-26B-A4B-it. The original model license and usage restrictions continue to apply.