CoolFace
Modelpublic

Yirasumi/jina-embeddings-v5-omni-nano-retrieval-GGUF

sourceHugging Facecc-by-nc-4.0updated 4mo agoView on Hugging Face
0likes47downloads
Model Card

jina-embeddings-v5-omni-nano-retrieval — GGUF

GGUF quantization of `jinaai/jina-embeddings-v5-omni-nano-retrieval` for use with llama.cpp.

The model produces 768-dim embeddings in a shared vector space for text, images, video, and audio (index with text, query with any modality).

⚠️ Important: Non-standard Q4KM

This Q4_K_M quant keeps the `token_embd` layer in F16 (--token-embedding-type f16).

A plain Q4KM (with quantized token embeddings) crashes on multimodal (image/video) inputs. Keeping the embedding layer in F16 fixes this while still shrinking the rest of the weights to Q4KM. Text-only quality is near-lossless vs F16 (cosine delta ~0.01–0.02).

FileSizeNotes
jina-embeddings-v5-omni-nano-retrieval-Q4_K_M.gguf~261 MBLLM weights, token_embd in F16
mmproj-jina-embeddings-v5-omni-nano-retrieval-F16.gguf~194 MBVision projector (required for image/video)

Requirements

You need the jina fork of llama.cpp (the feat-v5-omni branch), not upstream — upstream does not yet support this architecture:

bash
git clone --depth 1 --branch feat-v5-omni https://github.com/jina-ai/llama.cpp
cd llama.cpp && cmake -B build -DGGML_CUDA=ON && cmake --build build -j

Usage

Server (text + image)

bash
./build/bin/llama-server \
  -m jina-embeddings-v5-omni-nano-retrieval-Q4_K_M.gguf \
  --mmproj mmproj-jina-embeddings-v5-omni-nano-retrieval-F16.gguf \
  --embedding --pooling last \
  -c 8192 -b 8192 -ub 8192 \
  --host 127.0.0.1 --port 8080

Get the media marker (randomly generated per server start — required for image embedding):

bash
curl http://127.0.0.1:8080/props | jq -r .media_marker
# e.g. <__media_yAbTtTRgL15vbFiVDhX20zar2jGO88oM__>

Text embedding:

bash
curl -s http://127.0.0.1:8080/embeddings \
  -d '{"content":[{"prompt_string":"Query: Which planet is the Red Planet?"}]}'

Use Query: prefix for queries and Document: prefix for documents (retrieval-targeted model).

Image embedding (substitute the marker from /props):

bash
MARKER=$(curl -s http://127.0.0.1:8080/props | jq -r .media_marker)
IMG_B64=$(base64 -w0 photo.jpg)
curl -s http://127.0.0.1:8080/embeddings \
  -d "{\"content\":[{\"prompt_string\":\"$MARKER\",\"multimodal_data\":[\"$IMG_B64\"]}]}"

Text-only (any llama.cpp build)

For text-only embeddings, the standard llama-embedding works:

bash
./build/bin/llama-embedding \
  -m jina-embeddings-v5-omni-nano-retrieval-Q4_K_M.gguf \
  --pooling last --embd-normalize 2 \
  -p "Query: a cute cat" --embd-output-format json

Benchmarks (CPU, this quant)

Image embedding scales near-linearly with CPU cores (256×256 image, Intel Xeon E5-2683 v4):

ThreadsLatency
1~14.0 s
2~7.6 s
4~4.0 s

On GPU (e.g. T4 via PyTorch path) image embedding runs ~1.3 s. Text embedding is ~20–140 ms depending on backend.

Quality notes

  • —Cross-lingual same-meaning (EN↔CN): ~0.92 cosine
  • —Unrelated sentences: ~0.09 cosine
  • —Antonym pairs: ~0.70 (embedding models don't push antonyms toward −1; use an NLI model for that)
  • —Q4KM vs F16 text: cosine delta ~0.01–0.02

License

Inherits the base model license: CC-BY-NC-4.0 (non-commercial). See the original model card.