nlpai-lab/KURE-v2
<a href="https://github.com/nlpai-lab/KURE"> <img src="kure_logo.png" width="50%"/> </a>
π KURE-v2
KURE-v2 is a Korean-English bilingual late-interaction (multi-vector) retrieval model built on skt/A.X-Encoder-base. It encodes every token into a 128-dimensional vector and scores queryβdocument pairs with MaxSim.
At 154M parameters it averages nDCG@10 0.8160 across the nine MTEB(kor, v2) retrieval tasks β the strongest model on that suite, above every single-vector model measured, including one 175Γ its size.
It is trained in two stages: unsupervised contrastive pretraining on 20.7M unlabeled pairs (KURE-v2-unsupervised), then supervised fine-tuning with contrastive loss and KL distillation from a reranker.
Key Characteristics
- Late interaction: one 128-d vector per token, scored with MaxSim.
- Long documents: up to 8,192 tokens.
- Compact: 154M parameters.
- No instruction prefixes: queries and documents need no task instruction. Query expansion to 64 tokens is handled by the model.
Usage
<details> <summary><b>PyLate Usage</b></summary>
pip install -U pylateIndexing documents
from pylate import indexes, models, retrieve
model = models.ColBERT(model_name_or_path="nlpai-lab/KURE-v2")
index = indexes.PLAID(
index_folder="pylate-index",
index_name="index",
override=True,
)
documents_ids = ["1", "2", "3"]
documents = [
"μΈμ’
λμμ 1443λ
μ νλ―Όμ μμ μ°½μ νκ³ 1446λ
μ μ΄λ₯Ό λ°ν¬νμλ€.",
"κΉμΉλ λ°°μΆλ 무λ₯Ό μκΈμ μ μΈ λ€ κ³ μΆ§κ°λ£¨μ μ κ°μ λ£μ΄ λ°ν¨μν¨ μμμ΄λ€.",
"νλΌμ°μ ν΄λ° 1,947mλ‘ λ¨νμμ κ°μ₯ λμ μ°μ΄λ©° μ μ£Όλ μ€μμ μ리νλ€.",
]
documents_embeddings = model.encode(
documents,
batch_size=32,
is_query=False,
show_progress_bar=True,
)
index.add_documents(
documents_ids=documents_ids,
documents_embeddings=documents_embeddings,
)To reuse an existing index, instantiate it without override:
index = indexes.PLAID(index_folder="pylate-index", index_name="index")Retrieving top-k documents
retriever = retrieve.ColBERT(index=index)
queries_embeddings = model.encode(
["νλ―Όμ μμ μΈμ λ§λ€μ΄μ‘λμ?"],
batch_size=32,
is_query=True,
show_progress_bar=True,
)
scores = retriever.retrieve(
queries_embeddings=queries_embeddings,
k=10,
)Reranking
To rerank a first-stage candidate list without building an index:
from pylate import models, rank
model = models.ColBERT(model_name_or_path="nlpai-lab/KURE-v2")
queries = [
"μ κΈ°μ°¨ νλ°°ν°λ¦¬λ μ΄λ»κ² μ¬νμ©νλμ?",
"겨μΈμ νλΌμ°μ μ€λ₯Ό λ νμν μ₯λΉλ?",
]
documents = [
[
"νλ°°ν°λ¦¬μμ 리ν¬κ³Ό μ½λ°νΈλ₯Ό νμνλ μ΅μ μ λ ¨ 곡μ μ΄ μμ©νλκ³ μλ€.",
"κΈμ μΆ©μ κΈ°λ 30λΆ λ΄μΈλ‘ λ°°ν°λ¦¬λ₯Ό 80%κΉμ§ μΆ©μ ν μ μλ€.",
],
[
"κ²¨μΈ νλΌμ° μ°νμλ μμ΄μ κ³Ό λ°©ν μ₯κ°μ΄ νμμ΄λ©° μ
μ° μκ°μ΄ μ νλλ€.",
"μ μ£Ό μ¬λ κΈΈμ ν΄μμ λ°λΌ μ΄μ΄μ§λ 27κ° μ½μ€λ‘ ꡬμ±λμ΄ μλ€.",
"μ μ€κΈ°μλ λ±μ°νμ μ€ν¨μΈ λ₯Ό μ°©μ©ν΄ λμ΄ λ€μ΄κ°λ κ²μ λ§λ κ²μ΄ μ’λ€.",
],
]
documents_ids = [[1, 2], [1, 3, 2]]
queries_embeddings = model.encode(queries, is_query=True)
documents_embeddings = model.encode(documents, is_query=False)
reranked_documents = rank.rerank(
documents_ids=documents_ids,
queries_embeddings=queries_embeddings,
documents_embeddings=documents_embeddings,
)</details>
<details> <summary><b>Sentence-Transformers Usage</b></summary>
pip install -U sentence-transformersfrom sentence_transformers import MultiVectorEncoder
model = MultiVectorEncoder("nlpai-lab/KURE-v2")
query = "νλ―Όμ μμ μΈμ λ§λ€μ΄μ‘λμ?"
documents = [
"μΈμ’
λμμ 1443λ
μ νλ―Όμ μμ μ°½μ νκ³ 1446λ
μ μ΄λ₯Ό λ°ν¬νμλ€.",
"κΉμΉλ λ°°μΆλ 무λ₯Ό μκΈμ μ μΈ λ€ κ³ μΆ§κ°λ£¨μ μ κ°μ λ£μ΄ λ°ν¨μν¨ μμμ΄λ€.",
"νλΌμ°μ ν΄λ° 1,947mλ‘ λ¨νμμ κ°μ₯ λμ μ°μ΄λ©° μ μ£Όλ μ€μμ μ리νλ€.",
]
query_embeddings = model.encode_query(query)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings[0].shape)
# (64, 128) (29, 128)
# MaxSim late-interaction scoring (higher is more relevant)
scores = model.similarity(query_embeddings, document_embeddings)
print(scores)</details>
Evaluation
KURE-v2 is evaluated on the nine MTEB(kor, v2) Retrieval tasks. We report nDCG@10. The results are also shown in the official MTEB Leaderboard?types=Retrieval&s.summary=meanTask&d.summary=desc)
Late-interaction rows were measured with mteb 2.18.16 and PLAID retrieval. Single-vector rows are taken from the official MTEB results repository, except for Belebele, where only the Korean-query / Korean-corpus subset is used. The original version also includes cross-lingual subsets (Korean query β English corpus, English query β Korean corpus).
Serving
KURE-v2 is a late-interaction model: each document is stored as a set of token vectors, so the practical questions for deployment are index size and search cost. We benchmarked KURE-v2 across ANN backends and compression schemes on the 9 Korean MTEB retrieval tasks, against five single-vector baselines served with faiss HNSW. All numbers are end-to-end: batch-1 query encoding + index search, measured serially on one A100 80GB.
<p align="center"> <img src="assets/deploy_overview.png" width="100%" alt="Average nDCG@10 vs. index storage (left) and vs. end-to-end QPS (right)"> </p>
Two things the figures show:
- Hierarchical token pooling (x2) halves the index for a 0.04 nDCG drop. Asymmetric binary quantization (1-bit document tokens, bf16 queries) shrinks it 9.4x for 1.05. Stacking the two (pooling x3 + binary), the entire 9-corpus index fits in 1.7 GB, smaller than every single-vector HNSW index (13.1-50.0 GB), while still outscoring the best single-vector model (79.57 vs 79.07).
- A live query arrives as text: 4B-8B single-vector models spend 38-40 ms encoding it, capping them at ~25 QPS no matter how fast HNSW is. KURE-v2 encodes in 13.8 ms (154M params), so every configuration except MUVERA serves 43-55 QPS, roughly 2x the 8B single-vector models, at higher quality.
Large corpora: tail latency
<p align="center"> <img src="assets/bigcorpus_miracl.png" width="70%" alt="MIRACL (1.5M docs): quality, e2e p95 latency, index size"> </p>
On the largest corpus (MIRACL, ~1.5M documents) an exhaustive 1-bit scan costs O(corpus): p95 climbs to 156 ms, and pooling the tokens 3x only brings it to 74 ms. Generating candidates with faiss BinaryIVF (Hamming search over the same 1-bit index) and re-scoring them with exact asymmetric MaxSim cuts p95 to 38 ms on the same 2.2 GB index, lower tail latency than the 4B-8B single-vector baselines (43 ms) at higher nDCG. For large collections, use a candidate-generating index (PLAID or BinaryIVF), not an exhaustive scan.
<details> <summary><b>Measurement details</b></summary>
- Hardware: 1x NVIDIA A100 80GB, 2x AMD EPYC 7513 (64 cores), 1.2 TB RAM.
- Software: faiss-cpu 1.15.0, fast-plaid 1.6.0, sentence-transformers 6.0.0, PyTorch 2.8.0.
- Protocol: batch-1, serial. Index-search latency: 10 warmup queries, then every query of the task measured once (QPS = 1/mean). Query-encoding latency: 5 warmup, 50 measured. End-to-end = encoding + search.
- Precision: encoding in bf16; each index stores its own format (HNSW fp32, PLAID 4-bit residuals, binary 1-bit).
- Index size: the full serialized index on disk (vectors, graph, codebooks; external doc-id mapping excluded).
- Tasks: the 9 Korean MTEB retrieval tasks; MLDR is the mean of its dev/test splits; nDCG@10 x100.
- HNSW:
IndexHNSWFlat(inner product on L2-normalized embeddings), M=32, efConstruction=200, efSearch=64. - PLAID: nbits=4, all other settings fast-plaid defaults (kmeansniters=4, nivfprobe=8, nfull_scores=4096). nbits=2/1 give 27.0/17.0 GB at 81.25/81.09 nDCG.
- MUVERA: numrepetitions=10, numsimhashprojections=6, finalprojection_dimension=8192, exact-MaxSim rerank of the top 1,000.
- BinaryIVF: nlist=floor(sqrt(total tokens)) capped at 65,536, nprobe=32, top-128 Hamming tokens per query token, exact asymmetric-MaxSim rerank of the top 1,000 documents.
- Token pooling: hierarchical (Ward linkage), pool_factor 2-3, documents only. </details>
Citation
@misc{kure-v2,
title = {KURE-v2: a Korean-English bilingual late-interaction retriever},
author = {Jang, Youngjoon and Son, Junyoung and Lee, Taemin and Hong, Seongtae and Lim, Heuiseok},
year = {2026},
url = {https://huggingface.co/nlpai-lab/KURE-v2},
}@inproceedings{jang2025kure,
title={KURE: Embedding Model for Korean-Specific Retrieval},
author={Jang, Youngjoon and Son, Junyoung and Lee, Taemin and Hong, Seongtae and Park, JeongBae and Lim, Heuiseok},
booktitle={Annual Conference on Human and Language Technology},
pages={129--134},
year={2025},
organization={Human and Language Technology}
}@inproceedings{santhanam-etal-2022-colbertv2,
title = {ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction},
author = {Santhanam, Keshav and Khattab, Omar and Saad-Falcon, Jon and Potts, Christopher and Zaharia, Matei},
booktitle = {Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies},
year = {2022},
pages = {3715--3734},
}@misc{PyLate,
title = {PyLate: Flexible Training and Retrieval for Late Interaction Models},
author = {Chaffin, Antoine and Sourty, RaphaΓ«l},
year = {2024},
url = {https://github.com/lightonai/pylate},
}