CoolFace
Modelpublic

nanovdr/ColNanoVDR-Q-Ettin150M-Tomoro8B-320-ML

sourceHugging Facemitupdated 7d agoView on Hugging Face
1likes34downloads
Model Card

<h2 align="center">ColNanoVDR: Document-Free Query-Side Distillation for Multi-Vector Visual Document Retrieval</h2>

<p align="center"> <a href="https://huggingface.co/nanovdr">Models</a> &nbsp;|&nbsp; <a href="https://github.com/Ryenhails/NanoVDR">Code</a> &nbsp;|&nbsp; <a href="https://huggingface.co/datasets/nanovdr/NanoVDR-Train">Dataset</a> </p>


What is ColNanoVDR

ColNanoVDR extends NanoVDR from single-vector retrievers to the multi-vector, late-interaction encoders that define the state of the art in visual document retrieval (ColPali-style models such as ColQwen3.5, Vultron, Tomoro and ColVec).

In those systems the query is encoded, at serving time, by a multi-billion-parameter vision-language model, even though a query is plain text and carries no image. ColNanoVDR removes that model from the online path. Using document-free distillation, a small text-only encoder is trained to reproduce the teacher's query token embeddings under an optimal-transport objective, so its output lands directly in the teacher's late-interaction space and the teacher's existing page index can be scored by MaxSim unchanged. Training needs only cached teacher query embeddings: no page image is ever encoded, and no relevance label is ever used.

The result keeps 95.7% of a 8.8B teacher across 22 ViDoRe datasets with a query tower 59x smaller, and the document side is left untouched, so an existing index does not need to be rebuilt. With no vision tower on the query path, encoding a query becomes feasible on CPU, so serving no longer requires a GPU at all.

This model

Every ColNanoVDR release is named:

ColNanoVDR-Q-{backbone}-{teacher}-{dim}-ML
FieldMeaningThis model
Qthe query tower: the only side that is trained; documents stay with the teacherQ
{backbone}the text-only encoder distilled intoEttin150M = jhu-clsp/ettin-encoder-150m, 150M params
{teacher}the multi-vector VDR teacher whose embedding space this tower targetsTomoro8B = TomoroAI/tomoro-colqwen3-embed-8b, 8.8B params
{dim}late-interaction vector width of that teacher320
MLmultilingual training mixtureEnglish + 5 Latin-script European languages

So ColNanoVDR-Q-Ettin150M-Tomoro8B-320-ML is a 150M ettin-encoder-150m query tower that emits 320-d multi-vector queries inside the late-interaction space of TomoroAI/tomoro-colqwen3-embed-8b.

A tower is only valid with the teacher it was distilled from: the teacher and the width must both match. This one requires pages indexed by TomoroAI/tomoro-colqwen3-embed-8b; for a different teacher, take the corresponding tower:

Training

Dataset. nanovdr/NanoVDR-Train: 1.49M queries, 711K English plus 778K MarianMT translations into five Latin-script European languages. Only the query text is used. Teacher query embeddings are cached once before training, so the teacher never runs during training.

Loss: OTW (weighted entropic optimal transport). A query is a set of token vectors, and the student and the teacher do not produce the same number of tokens, so there is no token-to-token correspondence to regress onto. OTW instead treats each side as a distribution on the unit sphere, the student as mu_S = sum_i a_i d(s_i) and the teacher as mu_T = sum_j b_j d(t_j), and minimises the entropic optimal-transport cost between them under c(s,t) = 1 - <s,t>:

L = <P*, C>,   P* = argmin_{P in U(a,b)} <P, C> - eps H(P)

Transport is balanced, so every student token must carry mass and the student is forced to cover the teacher's full token distribution rather than collapsing onto a few easy directions. The W is the student marginal: a_i is a softmax over learned per-token weight logits instead of uniform, which lets a student token carry more mass and match a teacher measure with more atoms. The motivation is that the transport cost upper-bounds the MaxSim scoring error uniformly over every possible document, which is what makes the objective document-free: minimising it constrains retrieval behaviour without a single page ever being encoded.

Settings: eps = 0.05, 50 Sinkhorn iterations, AdamW one-cycle, peak LR 3e-4, 3% warmup, effective batch 512 (128 x 4 accumulation), 10 epochs.

Architecture. Ettin encoder -> bias-free linear projection to 320-d -> per-token L2 normalisation -> learned weight head. The weight head is a linear layer over the pre-projection hidden states whose softmax gives a_i; at inference each unit token vector is scaled by its weight, so the emitted vectors are deliberately not unit-length and the per-query token norms sum to 1. Special tokens are excluded from scoring. 150M parameters total.

Usage

Requires sentence-transformers>=6.0 and transformers>=5.0.

Retrieval runs in two stages. The teacher builds the page index once, offline; ColNanoVDR then answers every query online, and the teacher is never loaded again.

Step 1 (offline, once): index your pages with the teacher

This is the only stage where the vision-language model runs. Cache the resulting page embeddings; they are what you serve against.

python
import torch
from transformers import AutoModel, AutoProcessor

TEACHER = "TomoroAI/tomoro-colqwen3-embed-8b"
teacher = AutoModel.from_pretrained(
    TEACHER, trust_remote_code=True, dtype=torch.bfloat16,
    attn_implementation="sdpa", device_map="cuda:0"
).eval()
processor = AutoProcessor.from_pretrained(TEACHER, trust_remote_code=True)

page_embeddings = []
with torch.no_grad():
    for batch in batches_of(page_images, 4):          # PIL images
        inputs = processor.process_images(batch).to(teacher.device)
        page_embeddings.extend(teacher(**inputs))     # (n_patches, 320) each

# persist page_embeddings to disk / a multi-vector index (PLAID, Vespa, Qdrant, ...)

If you already run tomoro-colqwen3-embed-8b in production, skip this step entirely: your existing index is already in the right space and does not need rebuilding.

Step 2 (online, per query): encode with ColNanoVDR and score

python
from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder(
    "nanovdr/ColNanoVDR-Q-Ettin150M-Tomoro8B-320-ML",
    trust_remote_code=True,
)

q = model.encode_query(["What was the revenue growth in Q3 2024?"])
# list of (n_tokens, 320) float arrays, weights already folded in

scores = model.similarity(q, page_embeddings)   # meanMaxSim

No teacher, no image processor, no GPU required on this path.

Two notes that change results if ignored:

  • —Pass the raw query, with no instruction prefix. The teacher's targets were cached through the teacher's own query processor, but the student was trained to reproduce them from bare query text, and every number below was measured that way.
  • —Do not re-normalise the output. The learned weights are already folded in. similarity_fn_name is meanmaxsim (MaxSim over document tokens per query token, then averaged over query tokens); the summed ColBERT variant changes the ranking.

Performance

NDCG@5 on ViDoRe (22 datasets). The query side is measured in isolation against teacher-encoded pages, which attributes all error to this tower:

v1v2v3Avg
tomoro-colqwen3-embed-8b (teacher)90.6165.0059.0071.54
ColNanoVDR-Q-Ettin150M-Tomoro8B-320-ML89.9560.6154.8768.48
Retention99.3%93.3%93.0%95.7%

Efficiency

ParamsVision tower at query timeCPU serving
tomoro-colqwen3-embed-8b (teacher)8.8Brequiredno
ColNanoVDR-Q-Ettin150M-Tomoro8B-320-ML150Mnoneyes

The query tower is 59x smaller than the teacher and runs no vision tower, so a query can be encoded on CPU. Query latency depends on your hardware and on the teacher's attention kernels, so measure it on your own stack.

License

MIT. The teacher's own licence governs how you index and serve pages.

Contact

Zhuchenyang Liu, Aalto University: zhuchenyang.liu@aalto.fi

Citation

The ColNanoVDR paper is in preparation. For now, please cite NanoVDR:

bibtex
@article{nanovdr2026,
  title   = {NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M
             Text-Only Encoder for Visual Document Retrieval},
  author  = {Liu, Zhuchenyang and Zhang, Yao and Xiao, Yu},
  journal = {arXiv preprint arXiv:2603.12824},
  year    = {2026}
}