Yirasumi/jina-embeddings-v5-omni-nano-retrieval-GGUF
jina-embeddings-v5-omni-nano-retrieval — GGUF
GGUF quantization of `jinaai/jina-embeddings-v5-omni-nano-retrieval` for use with llama.cpp.
The model produces 768-dim embeddings in a shared vector space for text, images, video, and audio (index with text, query with any modality).
⚠️ Important: Non-standard Q4KM
This Q4_K_M quant keeps the `token_embd` layer in F16 (--token-embedding-type f16).
A plain Q4KM (with quantized token embeddings) crashes on multimodal (image/video) inputs. Keeping the embedding layer in F16 fixes this while still shrinking the rest of the weights to Q4KM. Text-only quality is near-lossless vs F16 (cosine delta ~0.01–0.02).
Requirements
You need the jina fork of llama.cpp (the feat-v5-omni branch), not upstream — upstream does not yet support this architecture:
git clone --depth 1 --branch feat-v5-omni https://github.com/jina-ai/llama.cpp
cd llama.cpp && cmake -B build -DGGML_CUDA=ON && cmake --build build -jUsage
Server (text + image)
./build/bin/llama-server \
-m jina-embeddings-v5-omni-nano-retrieval-Q4_K_M.gguf \
--mmproj mmproj-jina-embeddings-v5-omni-nano-retrieval-F16.gguf \
--embedding --pooling last \
-c 8192 -b 8192 -ub 8192 \
--host 127.0.0.1 --port 8080Get the media marker (randomly generated per server start — required for image embedding):
curl http://127.0.0.1:8080/props | jq -r .media_marker
# e.g. <__media_yAbTtTRgL15vbFiVDhX20zar2jGO88oM__>Text embedding:
curl -s http://127.0.0.1:8080/embeddings \
-d '{"content":[{"prompt_string":"Query: Which planet is the Red Planet?"}]}'Use Query: prefix for queries and Document: prefix for documents (retrieval-targeted model).
Image embedding (substitute the marker from /props):
MARKER=$(curl -s http://127.0.0.1:8080/props | jq -r .media_marker)
IMG_B64=$(base64 -w0 photo.jpg)
curl -s http://127.0.0.1:8080/embeddings \
-d "{\"content\":[{\"prompt_string\":\"$MARKER\",\"multimodal_data\":[\"$IMG_B64\"]}]}"Text-only (any llama.cpp build)
For text-only embeddings, the standard llama-embedding works:
./build/bin/llama-embedding \
-m jina-embeddings-v5-omni-nano-retrieval-Q4_K_M.gguf \
--pooling last --embd-normalize 2 \
-p "Query: a cute cat" --embd-output-format jsonBenchmarks (CPU, this quant)
Image embedding scales near-linearly with CPU cores (256×256 image, Intel Xeon E5-2683 v4):
On GPU (e.g. T4 via PyTorch path) image embedding runs ~1.3 s. Text embedding is ~20–140 ms depending on backend.
Quality notes
- Cross-lingual same-meaning (EN↔CN): ~0.92 cosine
- Unrelated sentences: ~0.09 cosine
- Antonym pairs: ~0.70 (embedding models don't push antonyms toward −1; use an NLI model for that)
- Q4KM vs F16 text: cosine delta ~0.01–0.02
License
Inherits the base model license: CC-BY-NC-4.0 (non-commercial). See the original model card.
