slovak-nlp/e5-sk-large
e5-sk-large
e5-sk-large is a Slovak text embedding model (365M parameters, 1024-dimensional embeddings) built by applying vocabulary trimming and fine-tuning to multilingual-e5-large. It achieves competitive performance with proprietary embedding APIs on SkMTEB — the first comprehensive Slovak text embedding benchmark — while being 35% smaller than the original model and fully locally deployable.
Released as part of the SkMTEB project (paper · GitHub · collection).
For a smaller, faster variant, see e5-sk-small (45M parameters).
Model Details
Usage
This model follows the standard E5 prefix convention: prepend query: to queries and passage: to documents during retrieval. For symmetric tasks (STS, clustering, classification), no prefix is needed.
With sentence-transformers
First install the Sentence Transformers library:
pip install -U sentence-transformersThen you can load this model and run inference.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("slovak-nlp/e5-sk-large")
# Retrieval
query_embedding = model.encode("query: Čo je hlavné mesto Slovenska?")
passage_embedding = model.encode("passage: Bratislava je hlavné a najväčšie mesto Slovenska.")
similarity = model.similarity(query_embedding, passage_embedding)
print(similarity) # tensor([[0.9269]])
# Batch encoding
sentences = [
"query: Aké je počasie v Bratislave?",
"passage: V Bratislave je dnes slnečno a teplo.",
"passage: Bratislava leží na brehu Dunaja.",
]
embeddings = model.encode(sentences)
print(embeddings.shape) # (3, 1024)With transformers directly
import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel
def average_pool(last_hidden_states, attention_mask):
last_hidden = last_hidden_states.masked_fill(~attention_mask[..., None].bool(), 0.0)
return last_hidden.sum(dim=1) / attention_mask.sum(dim=1)[..., None]
tokenizer = AutoTokenizer.from_pretrained("slovak-nlp/e5-sk-large")
model = AutoModel.from_pretrained("slovak-nlp/e5-sk-large")
texts = [
"query: Čo je hlavné mesto Slovenska?",
"passage: Bratislava je hlavné a najväčšie mesto Slovenska.",
]
batch_dict = tokenizer(texts, max_length=512, padding=True, truncation=True, return_tensors="pt")
with torch.no_grad():
outputs = model(**batch_dict)
embeddings = average_pool(outputs.last_hidden_state, batch_dict["attention_mask"])
embeddings = F.normalize(embeddings, p=2, dim=1)
print((embeddings[0] @ embeddings[1]).item())Prefix guide
Training
How it was built
Step 1 — Vocabulary Trimming. Before fine-tuning, Vocabulary Trimming (Ushio et al., 2023) was applied to multilingual-e5-large to remove tokens irrelevant to Slovak. Token frequencies were computed on FineWeb2-Slovak, a quality-filtered Slovak web corpus, and the top 60K tokens (out of 250K) were retained. This reduced the model from 560M to 365M parameters (35% reduction) without meaningful performance loss.
Step 2 — Fine-tuning. The trimmed model was fine-tuned on curated Slovak datasets from the skLEP benchmark:
Training configuration
Evaluation: SkMTEB Results
Evaluated on SkMTEB — 31 datasets across 7 task types. Scores are percentages (higher is better).
e5-sk-large is practically equivalent to `text-embedding-3-large` (TOST equivalence test: 90% CI within ±2 points), while being open-weight, locally deployable, and free to run. Cross-lingual Slovak–English and Slovak–Czech bitext mining performance is preserved within 1 F1 point compared to the original multilingual-e5-large.
<!---
Running the full SkMTEB evaluation
pip install mteb
mteb run -m slovak-nlp/e5-sk-large -b "MTEB(slk, v1)"--> ---
Intended Uses
- Semantic search and retrieval-augmented generation (RAG) over Slovak text
- Semantic textual similarity (STS)
- Text clustering and classification via embedding features
- Cross-lingual retrieval (Slovak–English, Slovak–Czech)
- Local deployment where API latency or cost is a concern
Limitations
- Optimised for Slovak; cross-lingual transfer to non-Slavic languages is not evaluated.
- Vocabulary trimming removes non-Slovak tokens; performance on heavily code-mixed text may be reduced.
- Training data skews toward news, parliamentary, and encyclopedic domains.
- Max sequence length during fine-tuning is 256 tokens (underlying architecture supports up to 512). ---
Citation
@inproceedings{suppa2025skmteb,
title = {{SkMTEB}: {Slovak} Massive Text Embedding Benchmark and Model Adaptation},
author = {{\v{S}}uppa, Marek and Ridzik, Andrej and Hl{\'a}dek, Daniel and
Kna{\v{z}}ekov{\'a}, Nat{\'a}lia and Ondrejov{\'a}, Vikt{\'o}ria},
year = {2025},
eprint = {2606.13647},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.13647}
}
@inproceedings{reimers-2019-sentence-bert,
title = {Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks},
author = {Reimers, Nils and Gurevych, Iryna},
booktitle = {Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing},
year = {2019},
publisher = {Association for Computational Linguistics},
url = {https://arxiv.org/abs/1908.10084}
}License
MIT
