CoolFace
Modelpublic

dahus/gemma-4-e2b-it-Q2_K-GGUF

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
2likes318downloads
Model Card

Gemma 4 E2B it — Q2_K GGUF

2-bit quantized GGUF version of google/gemma-4-e2b-it. Smallest and fastest variant in the series — use only if RAM is the hard constraint.

Other quantizations in this series: Q3_K_S · Q3_K_M · Q4_K_S · Q4_K_M · Q5_K_S · Q5_K_M · Q6_K · Q8

File Info

PropertyValue
FormatGGUF Q2_K
File size2.99 GB
Bits per weight~2
Size vs F163.1× smaller

Benchmark Results

Tested across 4 categories (Math, Logic, Code, Science), 3 prompts each. Greedy decoding, 200 max new tokens. Metrics compare logit distributions vs F16 baseline.

Results by Category

CategorySpeed (tok/s)SQNRTop-1 AgreementKL Divergence
🔢 Math30.95.0 dB35.3%3.8922
🧠 Logic31.75.7 dB33.8%4.1991
💻 Code31.86.7 dB24.4%4.4969
🔬 Science32.06.4 dB34.3%3.8713
Overall31.65.85 dB32.0%4.1149

Quantization Comparison

ModelSizeSpeed (tok/s)vs F16 speedSQNRTop-1 AgreeKL Div
F16 (baseline)8.67 GB5.71.0×baselinebaselinebaseline
Q2_K (this)2.99 GB31.65.6×5.85 dB32.0%4.1149
Q3KS3.11 GB28.95.1×10.12 dB63.2%1.2605
Q3KM2.98 GB27.44.8×13.93 dB63.2%1.6747
Q4KS3.37 GB25.04.4×19.10 dB80.9%0.3456
Q4KM3.43 GB24.04.2×20.33 dB82.4%0.3356
Q5KS3.6 GB21.93.9x23.32 dB87.7%0.1547
Q5KM3.63 GB22.03.9×23.25 dB86.9%0.1248
Q84.97 GB16.22.9×37.11 dB96.0%0.0171

Key Findings

  • —Quality: Significant degradation — only 32% Top-1 agreement with F16; output can be incoherent (see sample below)
  • —Speed: 31.6 tok/s — fastest in the series, 5.6× faster than F16
  • —Size: 2.78 GB — fits in under 4 GB RAM
  • —Best for: Extreme RAM-constrained environments where some output quality loss is acceptable; not recommended for reasoning or code tasks
⚠️ Warning: Q2K produces visibly broken outputs on this model. Sample response to a math prompt repeated token garbage (`skills skills skills...`). Consider Q3K_S or higher for usable results.

Usage

bash
# llama.cpp CLI
./llama-cli -m gemma-4-e2b-q2k.gguf -p "Solve step by step: 2 + 2 = ?" -n 200
python
# llama-cpp-python
from llama_cpp import Llama

llm = Llama(model_path="gemma-4-e2b-q2k.gguf", n_ctx=2048)
output = llm("Solve step by step: 2 + 2 = ?", max_tokens=200)
print(output["choices"][0]["text"])

Hardware

Tested on: CPU inference (llama.cpp) Context: 2048 tokens | Greedy decoding