ThakiCloud/SKILLRET-Edge-109M-int3-g32
SKILLRET-Edge-109M-int3-g32
A 109.5M-parameter bi-encoder for agent skill retrieval — picking the right skill out of a catalogue for a natural-language request. Distilled from ThakiCloud/SKILLRET-Embedding-0.6B and small enough to run on a CPU next to the agent.
Scored on the public ThakiCloud/SKILLRET test split (4,392 queries / 6,006 skills), the metric the SkillRet paper uses as its headline.
NDCG@10 77.10 · 54.7 MB packed
Results
Against fp16 this costs 2.08pp (77.10 vs 79.18, ±0.44), for a file 4.0x smaller (219.0 MB to 54.7 MB).
⛔ This model is specialised, and it does not generalise. The score above is on SkillRet, the task it was distilled for. On corpora it never saw it collapses: 5.81 on NFCorpus and 49.72 on SciFact, against roughly 31 to 39 and 64 to 78 for general-purpose embedders of similar or smaller size. Use it for skill retrieval; use a general embedder for anything else.
Usage
from sentence_transformers import SentenceTransformer
m = SentenceTransformer("ThakiCloud/SKILLRET-Edge-109M-int3-g32")
# ⛔ Encode queries BARE — no instruction prefix. See query_prefix.json.
q = m.encode(["I need to convert a spreadsheet into a chart"], normalize_embeddings=True)
d = m.encode(["Chart Builder — turn tabular data into bar/line charts"], normalize_embeddings=True)
print(q @ d.T)What "int3 / g32" means here
Weights are quantized group-wise, asymmetric min/max, group size 32, with scales and zero-points stored in fp16 — the same shape as a GGUF Q4K/Q80 block.
This repo ships two things, deliberately:
Be clear about what this does and does not buy you. The packed file is the real storage footprint and it is verified: predicted 54.924 MB vs 54.680 MB on disk (0.44% error). But the safetensors path runs fp32 arithmetic — the latency figures below were measured that way, and they are not integer-kernel speeds. Routing the packed weights through a real INT kernel (llama.cpp, ONNX Runtime) is unmeasured and is where further speedup would come from.
quantization.json carries the per-tensor layout if you want to write that kernel.
How it was trained
Knowledge distillation from ThakiCloud/SKILLRET-Embedding-0.6B (NDCG@10 78.48 on the same split), kd_weight=0.7, multi-positive InfoNCE over 1–3 golds per query, 12 epochs, cosine schedule. The epoch was chosen on a skill-disjoint holdout — the test split was never used for selection.
⚠️ Read this before you compare numbers
Use the same query prefix at train and eval time. This repo ships query_prefix.json recording the contract (resolved: "", i.e. queries are encoded bare, with no instruction prefix). Which prefix you pick barely matters after fine-tuning — none 79.18 vs an instruction prefix 78.50, inside the standard error — but keeping it consistent matters a great deal. We measured a +8.84pp swing on a single checkpoint purely from a train/eval prefix mismatch, and that mismatch also manufactured a fake result in which quantization appeared to beat fp16. It does not.
Other caveats worth stating plainly:
- The model card of the SkillRet reference models reports a 4,997-query / 6,660-skill split. That split is not in the currently published dataset — we verified the public files are hash-identical to ours at 4,392 / 6,006. Do not convert between the two.
- Standard error on this split is about ±0.45. Differences under ~1pp are not rankings.
- Pooling is CLS, embeddings are L2-normalized,
max_length=256at evaluation.
What did NOT work (so you don't repeat it)
License
Apache-2.0, inherited from the base model.
Citation
Benchmark: SkillRet, arXiv:2605.05726
Paper (measurements behind the quantized variants and the student frontier): Where Post-Training Quantization Breaks Text Embedders: A Measured Map Across Four Embedder Families, arXiv:2609.16391
@article{han2026ptqembedders,
title = {Where Post-Training Quantization Breaks Text Embedders: A Measured Map Across Four Embedder Families},
author = {Han, Hyojung},
journal = {arXiv preprint arXiv:2609.16391},
year = {2026}
}