jfan/gemma-4-E2B-it-webgpu-ONNX
Upload onnx/decoder_model_merged_q4f16.onnx with huggingface_hub
expose last_hidden_state (final-norm) for MTP drafter seeding
decoder: transposed q4f16 weights (MatMulNBits Pattern 2, byte-identical content)
decoder: rewrite DQ->MatMul to MatMulNBits Pattern 2 (DQ axis=0, 2D [K,N] weights)
Upload onnx/embed_tokens_q4f16.onnx_data_1 with huggingface_hub
Upload onnx/embed_tokens_q4f16.onnx with huggingface_hub
Upload config.json with huggingface_hub
Upload onnx/embed_tokens_q4f16.onnx_data_1 with huggingface_hub
Upload onnx/embed_tokens_q4f16.onnx_data with huggingface_hub
Upload onnx/embed_tokens_q4f16.onnx with huggingface_hub
WebGPU 4-bit QDQ export: decoder+embed+vision re-quantized from BF16 (all ops WebGPU-compatible), audio encoder carried over
initial commit
