CoolFace
Modelpublic

JoramMillenaar/paraphrase-multilingual-MiniLM-L12-v2-smol-variants

sourceHugging Faceapache-2.0updated 11d agoView on Hugging Face
0likes18downloads
Model Card

paraphrase-multilingual-MiniLM-L12-v2 — webgpu-safe

sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2, re-exported to ONNX with the embedding table and the encoder quantized independently, for use with Transformers.js.

Why this repo exists

The standard q8 dynamic-quantization build of this model produces wrong embeddings when run on ONNX Runtime Web's WebGPU backend. Two operators in that build — MatMulInteger and DynamicQuantizeLinear — are not correctly implemented on that execution provider as of ONNX Runtime Web 1.24. The failure is silent: no error is thrown, no NaN appears, the vectors are just wrong. Two unrelated sentences can come back with a cosine similarity above 0.9. (PR for ONNX Runtime fix is pending: https://github.com/microsoft/onnxruntime/pull/32579)

This repo quantizes only the embedding lookup table — the vocabulary matrix, which is 82% of this model's parameters — and leaves the encoder in full or half precision. The resulting graph never calls MatMulInteger or DynamicQuantizeLinear, so it is not affected by that bug on either backend.

Numbers below are from a real WebGPU run in Chrome, comparing each variant's cosine similarity on two unrelated sentences against the same pair embedded with the fp32 reference on CPU.

VariantSizeA·B (unrelated pair)Verdict
model_fp32.onnx470 MB−0.07correct
model_fp16.onnx235 MB−0.07correct
model_vocab8_fp32.onnx (this repo, WASM)182 MB−0.07correct
model_vocab8_fp32.onnx (this repo, WebGPU)182 MB0.94wrong — do not use
split model: int8 tables (JS) + model_encoder_fp32.onnx181 MB−0.07correct on both backends

The reference fp32 model at −0.07 on an unrelated pair is the expected result. Anything landing near 0 or below is behaving correctly; anything landing near 1 has collapsed.

Practical takeaway: the vocab8 ONNX files are safe on CPU (wasm) and unsafe on webgpu. If you need a single artifact that is correct on every backend with no capability detection, use the split model.

What's in this repo

onnx/
  model_fp32.onnx              full precision, every op supported everywhere
  model_fp16.onnx               half-precision encoder + vocabulary
  model_vocab8_fp32.onnx        int8 vocabulary, fp32 encoder — CPU only, see above
  model_vocab8_fp16.onnx        int8 vocabulary, fp16 encoder — CPU only, see above
  model_encoder_fp32.onnx       encoder only, starting after the embedding sum
  model_encoder_fp16.onnx       same, fp16
tables.bin                      int8 embedding tables (word/position/token-type), raw bytes
tables_q4.bin                   4-bit per-row embedding tables, raw bytes
split.json                      offsets, scales, and the cut-point tensor name for tables.bin / tables_q4.bin

model_encoder_*.onnx does not contain an embedding lookup. Its first input is the already-summed word + position + token-type embedding tensor (named in split.json as "cut"). You provide that tensor yourself by reading tables.bin (or tables_q4.bin) and doing the lookup in your own code — see below. This is what makes the split variant portable: the int8 tensor never enters the ONNX graph, so the runtime operator that mishandles it on WebGPU is never called.

Usage — standard dtype presets

For model_fp32.onnx / model_fp16.onnx / model_vocab8_fp32.onnx / model_vocab8_fp16.onnx, use Transformers.js normally:

js
import { pipeline } from '@huggingface/transformers';

const extractor = await pipeline(
  'feature-extraction',
  '<your-username>/paraphrase-multilingual-MiniLM-L12-v2-mixed-precision',
  {
    device: 'wasm',              // see the compatibility table above before using 'webgpu'
    model_file_name: 'model_vocab8_fp32',
  },
);

const output = await extractor('a sentence to embed', { pooling: 'mean', normalize: true });

Usage — split model (int8/int4 vocabulary + fp32/fp16 encoder)

This path runs correctly on WebGPU. It requires doing the embedding lookup yourself, in JavaScript, against tables.bin. Read split.json for the per-table byte offsets, scale, and zero point; each table is dequantized as (byte - zero_point) * scale. The three tables (word, token-type, position) are summed elementwise, then fed to model_encoder_fp32.onnx at the input named in split.json's "cut" field, alongside an attention_mask input.

position_offset in split.json matters: this model is built on XLM-RoBERTa, which numbers position ids starting at 2, not 0. Add it to your position index before the lookup.

A minimal reference implementation (tokenizer → lookup → encoder → mean pool) is in `build-variants.py` and the companion bench, `vocab-quant-compare.html`, which also measures speed, memory, and correctness across every variant in this repo and can be pointed at your own inputs.

How these were built

build-variants.py in this repo does the whole pipeline: fetches the base model, quantizes only the Gather ops with onnxruntime.quantization, converts to fp16 with onnxconverter_common where used, locates the embedding/encoder boundary automatically, and verifies every artifact actually loads and produces finite output before writing it. Run it yourself to reproduce this repo or to build the same variants for a different model:

bash
pip install onnx onnxruntime onnxconverter-common huggingface_hub
python build-variants.py --repo sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2

Known limitation

This does not fix MatMulInteger / DynamicQuantizeLinear on WebGPU — it avoids calling them. If your use case needs standard dynamic quantization (e.g. q8 from the original repo) on WebGPU, track microsoft/onnxruntime for that operator coverage; nothing in this repo is a substitute for that fix landing upstream.