CoolFace
Modelpublic

onnx-community/jina-embeddings-v4-vllm-retrieval

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes
Model Card

<br><br>

<p align="center"> <img src="https://huggingface.co/datasets/jinaai/documentation-images/resolve/main/logo.webp" alt="Jina AI: Your Search Foundation, Supercharged!" width="150px"> </p>

<p align="center"> <b>ONNX conversion of <a href="https://jina.ai/"><b>Jina AI</b></a>'s jina-embeddings-v4.</b> </p>

Jina Embeddings v4 → ONNX (sub-part decomposition)

Original Model | Blog | Technical Report | API

Model overview

Source checkpoints are the vLLM-merged, per-task variants of jina-embeddings-v4 — each is a stock Qwen2.5-VL-3B with one task LoRA merged into the base weights (no custom adapter code):

TaskSource repoPrompt prefixes
retrievaljinaai/jina-embeddings-v4-vllm-retrievalQuery: / Passage:
text-matchingjinaai/jina-embeddings-v4-vllm-text-matchingQuery:
codejinaai/jina-embeddings-v4-vllm-codeQuery:

These are single-vector models: the embedding is a masked mean-pool over the last hidden state, L2-normalized, 2048-d, with Matryoshka truncation to 128/256/512/1024/2048.

Decomposition

Rather than baking the ~6 GB LLM into one monolithic encoder per modality, the model is split into three ONNX sub-parts (same pattern as the other recipes in this repo). The heavy backbone is stored once and reused by both the text and image paths:

Sub-partInput → OutputNotes
vision.onnxpixel_values [N,1176] → image_features [N,2048]task-agnostic (vision tower has no LoRA); grid baked at build resolution
embeddings.onnxinput_ids [B,S] (+ image_features [N,2048]) → inputs_embeds [B,S,2048]token embeds; image features scattered into <image_pad> positions
backbone.onnxinputs_embeds, attention_mask, position_ids [3,B,S] → last_hidden [B,S,2048]shared LLM; MROPE position_ids host-computed

Compose at inference (all ONNX; the driver only wires sessions and pools):

text  : embeddings(ids)                     → backbone → mean-pool(attn_mask)   → L2norm
image : vision(px) → embeddings(ids, feats) → backbone → mean-pool(vision-span) → L2norm

Pooling and Matryoshka truncation happen in the driver (nothing baked), so one build serves every output dimension.

Why host-computed position_ids?

The model uses MROPE (mrope_section [16,24,24]), which onnxruntime-genai's ModelBuilder cannot emit. position_ids [3,B,S] are therefore computed on the host and fed in: cumulative positions for text, and get_rope_index(...) over the image grid for image inputs. The graph stays clean.

Files

Each build directory is self-contained (sub-parts + image_meta.npz + tokenizer/processor assets + manifest.json):

DirPrecisionvisionembeddingsbackbonetotal
cpu_fp16fp161.3 GB0.6 GB5.2 GB7.1 GB
cpu_fp32fp322.5 GB1.2 GB11 GB15 GB
cpu_int8fp16, backbone int81.3 GB0.6 GB2.8 GB4.6 GB

manifest.json records the source hf_model id (read from the checkpoint, not the folder name), precision, any quantized sub-parts, embedding dim, Matryoshka dims, and the composed flow.

cuda_* directories, if present and empty, are placeholders. This environment's PyTorch/ORT are CPU builds, so GPU builds produce nothing here. The exported ONNX is execution-provider agnostic — the same files run on CUDAExecutionProvider via onnxruntime-gpu with no rebuild and no device flag.

Fidelity vs full PyTorch

Composed ONNX chain vs the full Qwen2_5_VLForConditionalGeneration (pooled-embedding cosine, worst of 3 text samples + 1 image):

Buildworst cosineverdict
cpu_fp321.000000✅
cpu_fp160.999999✅
cpu_int8 (backbone)0.999601✅ (≥0.999)
int4 (backbone)0.921–0.954✅ (<0.94)

int8 is the quantization sweet spot — ~35 % smaller than fp16 with negligible cosine drift. int4 is too coarse for an embedding model (the pooled/normalized vector amplifies 4-bit weight error into 5–8 % drift, which wrecks ranking) and is intentionally not produced.

Reproducing / using

CPU only; runs in the repo's uv project env (transformers 5.x, torchvision for the image processor). The pipeline is three scripts sharing common.py:

bash
# build sub-parts — one --precision flag: fp16 (default) | fp32 | int8 | int4
uv run build.py --model vllm-retrieval --output onnx/cpu_fp16                  # fp16
uv run build.py --model vllm-retrieval --output onnx/cpu_fp32 --precision fp32
uv run build.py --model vllm-retrieval --output onnx/cpu_int8 --precision int8 # fp16 graph + int8 backbone
uv run build.py --model vllm-retrieval --output onnx/cpu_int4 --precision int4 # lossy (see below)

# eval — accepts multiple build dirs (positional), auto-detects each one's precision from its manifest
uv run eval.py --model vllm-retrieval onnx/cpu_fp16 onnx/cpu_fp32 onnx/cpu_int8

# inference (no PyTorch load)
uv run inference.py --onnx-dir onnx/cpu_fp16 --text "Overview of climate change impacts"
uv run inference.py --onnx-dir onnx/cpu_fp16 --text "..." --prefix Passage --truncate-dim 256
uv run inference.py --onnx-dir onnx/cpu_fp16 --image doc.png

int8/int4 build the fp16 graph, then weight-quantize the backbone in place (block-wise MatMulNBits; vision/embeddings stay fp16). build.py runs a composed self-sanity check: fp16 / fp32 / int8 must hit cosine ≥ 0.999 or the build fails, while int4 only warns (it's knowingly lossy — verify with eval.py). Point --model at vllm-retrieval, vllm-text-matching, or vllm-text-code to build the other tasks. The vision sub-part is identical across tasks (no LoRA), so a vision.onnx can be shared to save disk.

Minimal ONNX Runtime example (text)

python
import json, numpy as np, onnxruntime as ort
from pathlib import Path
from transformers import AutoTokenizer

d = Path("onnx/cpu_fp16")
man = json.loads((d / "manifest.json").read_text())
npdt = np.float16 if man["precision"] == "fp16" else np.float32
tok = AutoTokenizer.from_pretrained(str(d))

def sess(name):  # log level raised to silence the harmless constant-fold notice
    so = ort.SessionOptions(); so.log_severity_level = 3
    return ort.InferenceSession(str(d / name), so, providers=["CPUExecutionProvider"])

emb_s, back_s = sess("embeddings.onnx"), sess("backbone.onnx")

enc = tok(["Query: Overview of climate change impacts"], return_tensors="np", padding="longest")
ids, am = enc["input_ids"], enc["attention_mask"]
pos = np.clip(np.cumsum(am, -1) - 1, 0, None)[None].repeat(3, 0)         # MROPE (text)
e = emb_s.run(None, {"input_ids": ids, "image_features": np.zeros((0, 2048), npdt)})[0]
h = back_s.run(None, {"inputs_embeds": e, "attention_mask": am, "position_ids": pos})[0]

pooled = (h * am[..., None]).sum(1) / am.sum(1, keepdims=True)           # masked mean-pool
emb = pooled / np.linalg.norm(pooled, axis=-1, keepdims=True)           # L2-norm → [1, 2048]