CoolFace
20 results

gemma-4-e2b

juiceb0xc0de /gemma-4-e2b-atlas image1M<n<10M4 likes839 downloads7d agoHugging FaceJWei05 /gemma4-e2b-base-topk128-traces0 likes425 downloads2mo agoHugging FaceJWei05 /gemma4-e2b-base-topk128-hf-overlay-v128-seed42 Gemma 4 E2B base top-k-128 HF training overlay This is the immutable training-engine overlay used to distill traces from Gemma 4 E2B base into Gemma 4 E4B. It preserves the prompts, responses, and exact response token IDs from JWei05/gemma4-e2b-base-topk128-traces, but replaces the source vLLM top-k targets with targets recomputed by the Hugging Face training engine. This repository is a reproducibility artifact for the corresponding distillation run. It is not a new… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/gemma4-e2b-base-topk128-hf-overlay-v128-seed42.tabulartext-generation10K<n<100K0 likes309 downloads2mo agoHugging Facejuiceb0xc0de /gemma-4-e2b-it-SAE gemma-4-e2b-it — 35-layer SAE atlas Sparse autoencoders on every decoder layer of gemma-4-e2b-it. Trained from scratch in one rolling pipeline with an event-aware controller. 35 layers, 49,152 features per layer, no per-layer hand-tuning. The base model is a stubborn one. 15 sliding-window layers, then BAM no KV cache, thick and getting thicker the deeper you go. This atlas was built the whole way through it anyway. What this is Three months of work. My first… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/gemma-4-e2b-it-SAE.feature-extraction10B<n<100B4 likes273 downloads28d agoHugging FaceChengshuo0723 /gemma-4-E2B-it-pi-mono-agent-loss-fp0 likes157 downloads2mo agoHugging FaceChengshuo0723 /gemma-4-E2B-it-pi-mono-agent-eval0 likes105 downloads2mo agoHugging Face