CoolFace
Modelpublic

jfan/gemma-4-E2B-it-webgpu-ONNX

sourceHugging Faceupdated 16d agoView on Hugging Face
0likes114downloads
12 commits on main
75eefe616d ago

Upload onnx/decoder_model_merged_q4f16.onnx with huggingface_hub

jfan
4f3ad2416d ago

expose last_hidden_state (final-norm) for MTP drafter seeding

jfan
a3001e516d ago

decoder: transposed q4f16 weights (MatMulNBits Pattern 2, byte-identical content)

jfan
3b764e416d ago

decoder: rewrite DQ->MatMul to MatMulNBits Pattern 2 (DQ axis=0, 2D [K,N] weights)

jfan
385a97317d ago

Upload onnx/embed_tokens_q4f16.onnx_data_1 with huggingface_hub

jfan
782064f17d ago

Upload onnx/embed_tokens_q4f16.onnx with huggingface_hub

jfan
444d5ca17d ago

Upload config.json with huggingface_hub

jfan
4dd5ed917d ago

Upload onnx/embed_tokens_q4f16.onnx_data_1 with huggingface_hub

jfan
6f530e217d ago

Upload onnx/embed_tokens_q4f16.onnx_data with huggingface_hub

jfan
02d491417d ago

Upload onnx/embed_tokens_q4f16.onnx with huggingface_hub

jfan
48fa11f17d ago

WebGPU 4-bit QDQ export: decoder+embed+vision re-quantized from BF16 (all ops WebGPU-compatible), audio encoder carried over

jfan
f2ad08017d ago

initial commit

jfan