CoolFace
Modelpublic

raghunath1/Aptivra-Base-110M-GGUF

sourceHugging Facemitupdated 4mo agoView on Hugging Face
1likes29downloads
Model Card

Aptivra-Base-110M — GGUF

GGUF (llama.cpp) builds of `raghunath1/Aptivra-Base-110M`, a 110M-parameter sentence-embedding model for skill routing and semantic retrieval.

This is an embedding model, not a chat model.

✅ Use it for❌ Do not use it for
query / document embeddingschat completion
skill routinginstruction following
semantic retrievaltext generation
vector search(it produces vectors, not text)
candidate ranking
⚠️ Experimental research preview — not production-certified. Validate on your own cases. These files are the embedding model only (no reranker). Full evidence, disclaimer, and the PyTorch/ONNX/OpenVINO builds are in the main repo `raghunath1/Aptivra-Base-110M`.

Files

FileQuantSizeFidelity vs fp32Notes
Aptivra-Base-110M-F16.ggufF16209 MB0.99999full precision; matches PyTorch, Recall@1 ≡ reference
Aptivra-Base-110M-Q8_0.ggufQ8_0112 MB0.99984near-lossless; recommended default, Recall@1 ≡ reference
Aptivra-Base-110M-Q4_K_M.ggufQ4KM71 MB0.98618smallest; ~1.4% perturbation, Recall@1 ≈ reference

Parity

Fidelity = mean cosine of each quant's embeddings to the fp32 reference, on the identical plain-text routing eval. F16/Q80 are effectively lossless (Recall@1 equals the [fp32 reference](https://huggingface.co/raghunath1/Aptivra-Base-110M)'s **0.958**); Q4KM trades ~1.4% embedding fidelity for the smallest footprint (71 MB). Pick Q80 unless you need the size.

Input format (important)

Feed plain text — no `query:` / `passage:` prefix (this fine-tune was trained without them). Use mean pooling and L2-normalized embeddings; compare with cosine similarity.

Usage — llama.cpp embedding mode

Build/run with mean pooling + L2 normalize (--pooling mean --embd-normalize 2):

bash
llama-embedding -m Aptivra-Base-110M-Q8_0.gguf \
  -p "set up a browser automation task" \
  --pooling mean --embd-normalize 2

Server (OpenAI-compatible embeddings endpoint):

bash
llama-server -m Aptivra-Base-110M-Q8_0.gguf --embeddings --pooling mean
# then: curl http://localhost:8080/v1/embeddings -d '{"input":"semantic search query"}'

Install llama.cpp per OS

  • —macOS: brew install llama.cpp
  • —Linux: brew install llama.cpp, or build from source (cmake -B build && cmake --build build), or use the prebuilt release binaries from the llama.cpp releases page.
  • —Windows: winget install llama.cpp, or download the prebuilt llama-*-bin-win-*.zip from the llama.cpp releases page (CUDA/Vulkan/CPU variants available).

Download a single file

bash
pip install huggingface_hub
huggingface-cli download raghunath1/Aptivra-Base-110M-GGUF \
  Aptivra-Base-110M-Q8_0.gguf --local-dir .
LM Studio / Ollama caveat: both are built around chat/completion models. This is an embedding model — use it only through an embeddings path (llama.cpp llama-server /v1/embeddings, or Ollama's /api/embeddings), never the chat UI. It returns vectors, not text.

Provenance

Converted from the canonical safetensors with llama.cpp convert_hf_to_gguf.py (F16), then llama-quantize for Q8_0 and Q4_K_M. Each quant is gated by the Recall@1 parity table above before release. Architecture: BERT (e5-base-v2), 768-dim, ctx 512, mean pooling.

License

MIT. Fine-tuned from intfloat/e5-base-v2 (MIT).