hexoy/gemma-4-e2b-distilled
Gemma 4 E2B Distilled
Distilled Gemma 4 E2B with all 35 language-model MLPs replaced by two-factor Monarch maps. The released BF16 model has 3,682,268,704 parameters, instead of 5,104,297,504 in the original Gemma 4 E2B.
1. Weight Storage
BF16 sizes use two bytes per parameter; serialized LoRA size was audited from the two safetensor shards. Values exclude activations and CUDA workspaces.
2. TinyHellaSwag
All rows used the same RTX PRO 6000, fixed batch size 32, seed 1234, 100 official examples, 10-shot prompts, and no chat template.
Rank-8 LoRA improved the 35-layer source by 0.90 GP-IRT points and one correct item, but missed the former +1.0 point or +2 item release gate. The 100-example benchmark is noisy.
Usage
from transformers import AutoModelForImageTextToText, AutoProcessor
model_id = "hexoy/gemma-4-e2b-distilled"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
model_id,
trust_remote_code=True,
dtype="auto",
device_map="auto",
)The rank-8 result is published separately at `hexoy/gemma-4-e2b-monarch-35mlp-lora-r8`. It is a benchmark-recovery experiment: corrected 64-token generation repeated phrases for text and remained malformed or inaccurate for images. It is not a general text or multimodal recovery release.
Reproducibility
- BF16 model revision:
f897353fca328b1cc5fd2e12d645773ca637f5f0 - GitHub repository: `ratmir-miftachov/gemma-distillation`
- INT8 variant: `hexoy/gemma-4-e2b-monarch-35mlp-int8`
- Experimental LoRA r8 variant: `hexoy/gemma-4-e2b-monarch-35mlp-lora-r8`
- LoRA source commit:
1435571b20dd26c073a535678975884154add5b8
Detailed training and release evidence is retained privately.
Derived from google/gemma-4-E2B-it. See NOTICE for the modification summary.
