gemma-4-e2b
gemma-4-e2b-atlas
gemma4-e2b-base-topk128-tracesgemma4-e2b-base-topk128-hf-overlay-v128-seed42
Gemma 4 E2B base top-k-128 HF training overlay
This is the immutable training-engine overlay used to distill traces from Gemma 4 E2B base into
Gemma 4 E4B. It preserves the prompts, responses, and exact response token IDs from
JWei05/gemma4-e2b-base-topk128-traces,
but replaces the source vLLM top-k targets with targets recomputed by the Hugging Face training
engine.
This repository is a reproducibility artifact for the corresponding distillation run. It is not a
new… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/gemma4-e2b-base-topk128-hf-overlay-v128-seed42.gemma-4-e2b-it-SAE
gemma-4-e2b-it — 35-layer SAE atlas
Sparse autoencoders on every decoder layer of gemma-4-e2b-it. Trained from scratch in one rolling pipeline with an event-aware controller. 35 layers, 49,152 features per layer, no per-layer hand-tuning.
The base model is a stubborn one. 15 sliding-window layers, then BAM no KV cache, thick and getting thicker the deeper you go. This atlas was built the whole way through it anyway.
What this is
Three months of work. My first… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/gemma-4-e2b-it-SAE.gemma-4-E2B-it-pi-mono-agent-loss-fpgemma-4-E2B-it-pi-mono-agent-eval
