raghunath1/Aptivra-Base-110M-GGUF
Aptivra-Base-110M — GGUF
GGUF (llama.cpp) builds of `raghunath1/Aptivra-Base-110M`, a 110M-parameter sentence-embedding model for skill routing and semantic retrieval.
This is an embedding model, not a chat model.
⚠️ Experimental research preview — not production-certified. Validate on your own cases. These files are the embedding model only (no reranker). Full evidence, disclaimer, and the PyTorch/ONNX/OpenVINO builds are in the main repo `raghunath1/Aptivra-Base-110M`.
Files
Parity
Fidelity = mean cosine of each quant's embeddings to the fp32 reference, on the identical plain-text routing eval. F16/Q80 are effectively lossless (Recall@1 equals the [fp32 reference](https://huggingface.co/raghunath1/Aptivra-Base-110M)'s **0.958**); Q4KM trades ~1.4% embedding fidelity for the smallest footprint (71 MB). Pick Q80 unless you need the size.
Input format (important)
Feed plain text — no `query:` / `passage:` prefix (this fine-tune was trained without them). Use mean pooling and L2-normalized embeddings; compare with cosine similarity.
Usage — llama.cpp embedding mode
Build/run with mean pooling + L2 normalize (--pooling mean --embd-normalize 2):
llama-embedding -m Aptivra-Base-110M-Q8_0.gguf \
-p "set up a browser automation task" \
--pooling mean --embd-normalize 2Server (OpenAI-compatible embeddings endpoint):
llama-server -m Aptivra-Base-110M-Q8_0.gguf --embeddings --pooling mean
# then: curl http://localhost:8080/v1/embeddings -d '{"input":"semantic search query"}'Install llama.cpp per OS
- macOS:
brew install llama.cpp - Linux:
brew install llama.cpp, or build from source (cmake -B build && cmake --build build), or use the prebuilt release binaries from the llama.cpp releases page. - Windows:
winget install llama.cpp, or download the prebuiltllama-*-bin-win-*.zipfrom the llama.cpp releases page (CUDA/Vulkan/CPU variants available).
Download a single file
pip install huggingface_hub
huggingface-cli download raghunath1/Aptivra-Base-110M-GGUF \
Aptivra-Base-110M-Q8_0.gguf --local-dir .LM Studio / Ollama caveat: both are built around chat/completion models. This is an embedding model — use it only through an embeddings path (llama.cppllama-server/v1/embeddings, or Ollama's/api/embeddings), never the chat UI. It returns vectors, not text.
Provenance
Converted from the canonical safetensors with llama.cpp convert_hf_to_gguf.py (F16), then llama-quantize for Q8_0 and Q4_K_M. Each quant is gated by the Recall@1 parity table above before release. Architecture: BERT (e5-base-v2), 768-dim, ctx 512, mean pooling.
License
MIT. Fine-tuned from intfloat/e5-base-v2 (MIT).
