CoolFace
Modelpublic

salohcin714/gemma-4-E2B-it-5bit-mlx

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes25downloads
Model Card

gemma-4-E2B-it-5bit-mlx

Provenance

Converted from `google/gemma-4-E2B-it` using mlx-lm 0.31.3.

Usage

python
from mlx_lm import load, generate

model, tokenizer = load("salohcin714/gemma-4-E2B-it-5bit-mlx")
messages = [{"role": "user", "content": "Hello"}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
text = generate(model, tokenizer, prompt=prompt, verbose=True)

Modifications

Weights converted to MLX safetensors layout and quantized (5-bit affine quantization, group size 64, via round-to-nearest, no calibration). No fine-tuning; no added training data.

License and attribution

Licensed under Apache 2.0. Original weights by Google. See the upstream model card and the included LICENSE file for the full text.

Disclaimer

This repository is not affiliated with or endorsed by Google. "Gemma" is a Google trademark, used here descriptively to identify the origin of the base model.