JoramMillenaar/paraphrase-multilingual-MiniLM-L12-v2-smol-variants
paraphrase-multilingual-MiniLM-L12-v2 — webgpu-safe
sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2, re-exported to ONNX with the embedding table and the encoder quantized independently, for use with Transformers.js.
Why this repo exists
The standard q8 dynamic-quantization build of this model produces wrong embeddings when run on ONNX Runtime Web's WebGPU backend. Two operators in that build — MatMulInteger and DynamicQuantizeLinear — are not correctly implemented on that execution provider as of ONNX Runtime Web 1.24. The failure is silent: no error is thrown, no NaN appears, the vectors are just wrong. Two unrelated sentences can come back with a cosine similarity above 0.9. (PR for ONNX Runtime fix is pending: https://github.com/microsoft/onnxruntime/pull/32579)
This repo quantizes only the embedding lookup table — the vocabulary matrix, which is 82% of this model's parameters — and leaves the encoder in full or half precision. The resulting graph never calls MatMulInteger or DynamicQuantizeLinear, so it is not affected by that bug on either backend.
Numbers below are from a real WebGPU run in Chrome, comparing each variant's cosine similarity on two unrelated sentences against the same pair embedded with the fp32 reference on CPU.
The reference fp32 model at −0.07 on an unrelated pair is the expected result. Anything landing near 0 or below is behaving correctly; anything landing near 1 has collapsed.
Practical takeaway: the vocab8 ONNX files are safe on CPU (wasm) and unsafe on webgpu. If you need a single artifact that is correct on every backend with no capability detection, use the split model.
What's in this repo
onnx/
model_fp32.onnx full precision, every op supported everywhere
model_fp16.onnx half-precision encoder + vocabulary
model_vocab8_fp32.onnx int8 vocabulary, fp32 encoder — CPU only, see above
model_vocab8_fp16.onnx int8 vocabulary, fp16 encoder — CPU only, see above
model_encoder_fp32.onnx encoder only, starting after the embedding sum
model_encoder_fp16.onnx same, fp16
tables.bin int8 embedding tables (word/position/token-type), raw bytes
tables_q4.bin 4-bit per-row embedding tables, raw bytes
split.json offsets, scales, and the cut-point tensor name for tables.bin / tables_q4.binmodel_encoder_*.onnx does not contain an embedding lookup. Its first input is the already-summed word + position + token-type embedding tensor (named in split.json as "cut"). You provide that tensor yourself by reading tables.bin (or tables_q4.bin) and doing the lookup in your own code — see below. This is what makes the split variant portable: the int8 tensor never enters the ONNX graph, so the runtime operator that mishandles it on WebGPU is never called.
Usage — standard dtype presets
For model_fp32.onnx / model_fp16.onnx / model_vocab8_fp32.onnx / model_vocab8_fp16.onnx, use Transformers.js normally:
import { pipeline } from '@huggingface/transformers';
const extractor = await pipeline(
'feature-extraction',
'<your-username>/paraphrase-multilingual-MiniLM-L12-v2-mixed-precision',
{
device: 'wasm', // see the compatibility table above before using 'webgpu'
model_file_name: 'model_vocab8_fp32',
},
);
const output = await extractor('a sentence to embed', { pooling: 'mean', normalize: true });Usage — split model (int8/int4 vocabulary + fp32/fp16 encoder)
This path runs correctly on WebGPU. It requires doing the embedding lookup yourself, in JavaScript, against tables.bin. Read split.json for the per-table byte offsets, scale, and zero point; each table is dequantized as (byte - zero_point) * scale. The three tables (word, token-type, position) are summed elementwise, then fed to model_encoder_fp32.onnx at the input named in split.json's "cut" field, alongside an attention_mask input.
position_offset in split.json matters: this model is built on XLM-RoBERTa, which numbers position ids starting at 2, not 0. Add it to your position index before the lookup.
A minimal reference implementation (tokenizer → lookup → encoder → mean pool) is in `build-variants.py` and the companion bench, `vocab-quant-compare.html`, which also measures speed, memory, and correctness across every variant in this repo and can be pointed at your own inputs.
How these were built
build-variants.py in this repo does the whole pipeline: fetches the base model, quantizes only the Gather ops with onnxruntime.quantization, converts to fp16 with onnxconverter_common where used, locates the embedding/encoder boundary automatically, and verifies every artifact actually loads and produces finite output before writing it. Run it yourself to reproduce this repo or to build the same variants for a different model:
pip install onnx onnxruntime onnxconverter-common huggingface_hub
python build-variants.py --repo sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2Known limitation
This does not fix MatMulInteger / DynamicQuantizeLinear on WebGPU — it avoids calling them. If your use case needs standard dynamic quantization (e.g. q8 from the original repo) on WebGPU, track microsoft/onnxruntime for that operator coverage; nothing in this repo is a substitute for that fix landing upstream.
