CoolFace
Modelpublic

bokuweb/gemma-4-E4B-it-grande-wgpu-ja

sourceHugging Facegemmaupdated 6d agoView on Hugging Face
0likes73downloads
Model Card

bokuweb/gemma-4-E4B-it-grande-wgpu-ja

Gemma 4 E4B (Q40 codes from the llama.cpp GGUF, repacked) for [grande](https://github.com/bokuweb/grande)'s wgpu engine: state + every question in one block-causal forward pass, in the browser on WebGPU or natively on Metal / Vulkan. Exported with `tools/exportwgpu_gguf.py; manifest.json` maps tensors to files. Not a standalone checkpoint format.

Vocabulary pruned to the 25,392 tokens a Japanese / English decision corpus uses (tools/prune_vocab.py); text inside that corpus tokenizes exactly as with the full vocabulary, other text into somewhat more pieces.

bash
grande probe --model <this directory> --request examples/ticket-ja.json