CoolFace
Modelpublic

positron-ai/google_gemma-3-27b-it-ingest-best-gptq

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes555downloads
Model Card

Positron AI Quantized Build

This repository contains a Positron AI quantized build of google/gemma-3-27b-it for inference.

Recommended Use

Use this artifact when you need a GPTQ 4-bit build of google/gemma-3-27b-it built by Positron AI.

For general-purpose GPU inference, compare against the original model and other quantized formats before deployment.

Artifact Summary

FieldValue
Base modelgoogle/gemma-3-27b-it
Published artifactpositron-ai/google_gemma-3-27b-it-ingest-best-gptq
Quantization methodGPTQ
Quantization formatgptq
Source precisionn/a
Target runtimen/a
Hardware targetn/a
Release date2026-08-25
Licenseother

Quantization Details

FieldValue
Weight precision4-bit
Activation precisionnot quantized
Bits4
Group size64
Symmetric quantizationtrue
Activation ordering / desc_actfalse
Damp percent0.05
Calibration datasetMixed-domain calibration set
Calibration samples128
Calibration sequence length4096
MoE experts per tokenn/a
Quantization toolchainGPTQModel 7.2.0, transformers 5.11.0, torch 2.9.1, CUDA 12.8

Validation Results

MetricResultReferenceNotes
Mean KL-divergencen/an/aNot measured for this release
P95 KL-divergencen/an/aNot measured for this release
Top-1 agreementn/an/aNot measured for this release
Perplexity / NLL deltan/an/aNot measured for this release
MMLU meanpendingn/aEvaluation pending

KL-divergence has not been measured for this release; the metrics above will be populated when validation completes.

Evaluation Methodology

FieldValue
Evaluation daten/a
Evaluation suiten/a
Number of promptsn/a
Runtimen/a
Devicen/a
Pass criterian/a

Known Limitations

  • —MMLU evaluation is pending; results will be added when available.

Provenance

This artifact was produced by Positron AI from google/gemma-3-27b-it. The original model license and usage restrictions continue to apply.