ThakiCloud/SKILLRET-Edge-109M
SKILLRET-Edge-109M
A 109.5M-parameter bi-encoder for agent skill retrieval — picking the right skill out of a catalogue for a natural-language request. Distilled from ThakiCloud/SKILLRET-Embedding-0.6B and small enough to run on a CPU next to the agent.
Scored on the public ThakiCloud/SKILLRET test split (4,392 queries / 6,006 skills), the metric the SkillRet paper uses as its headline.
NDCG@10 79.18 · 219 MB fp16
Correction — September 2026. Earlier versions of this card reported the SKILLRET-Embedding-0.6B teacher as 80.82 NDCG@10 on the current 4,392-query / 6,006-skill evaluation split. That score was measured before the train/evaluation query-prefix contract was corrected. Re-evaluation onThakiCloud/SKILLRETrevisiona050ad2under the canonical query contract gives 78.48 NDCG@10. The model weights are unchanged. The earlier 605 MB size was also not the serialized artifact size; the distributed BF16 checkpoint is 1191.6 MB (decimal MB). We retain this note so results quoted from earlier versions of the card can be interpreted correctly.
Results
This model recovers 98% of a 26x larger teacher. int8 and int4 are free; int3 costs 1.14pp and is the point where quantization starts to be a real trade.
Usage
from sentence_transformers import SentenceTransformer
m = SentenceTransformer("ThakiCloud/SKILLRET-Edge-109M")
# ⛔ Encode queries BARE — no instruction prefix. See query_prefix.json.
q = m.encode(["I need to convert a spreadsheet into a chart"], normalize_embeddings=True)
d = m.encode(["Chart Builder — turn tabular data into bar/line charts"], normalize_embeddings=True)
print(q @ d.T)How it was trained
Knowledge distillation from ThakiCloud/SKILLRET-Embedding-0.6B (NDCG@10 78.48 on the same split), kd_weight=0.7, multi-positive InfoNCE over 1–3 golds per query, 12 epochs, cosine schedule. The epoch was chosen on a skill-disjoint holdout — the test split was never used for selection.
⚠️ Read this before you compare numbers
Use the same query prefix at train and eval time. This repo ships query_prefix.json recording the contract (resolved: "", i.e. queries are encoded bare, with no instruction prefix). Which prefix you pick barely matters after fine-tuning — none 79.18 vs an instruction prefix 78.50, inside the standard error — but keeping it consistent matters a great deal. We measured a +8.84pp swing on a single checkpoint purely from a train/eval prefix mismatch, and that mismatch also manufactured a fake result in which quantization appeared to beat fp16. It does not.
Other caveats worth stating plainly:
- The model card of the SkillRet reference models reports a 4,997-query / 6,660-skill split. That split is not in the currently published dataset — we verified the public files are hash-identical to ours at 4,392 / 6,006. Do not convert between the two.
- Standard error on this split is about ±0.45. Differences under ~1pp are not rankings.
- Pooling is CLS, embeddings are L2-normalized,
max_length=256at evaluation.
What did NOT work (so you don't repeat it)
License
Apache-2.0, inherited from the base model.
Citation
Benchmark: SkillRet, arXiv:2605.05726
Paper (measurements behind the quantized variants and the student frontier): Where Post-Training Quantization Breaks Text Embedders: A Measured Map Across Four Embedder Families, arXiv:2609.16391
@article{han2026ptqembedders,
title = {Where Post-Training Quantization Breaks Text Embedders: A Measured Map Across Four Embedder Families},
author = {Han, Hyojung},
journal = {arXiv preprint arXiv:2609.16391},
year = {2026}
}