CoolFace
Datasetpublic

robro612/vidore3_computerscience_neomme_260m_li

vidore3_computerscience_neomme_260m_li Multi-vector (late-interaction) embeddings of ViDoRe computerscience (vidore/computerscience), encoded with Hcompany/NeoMME-260M-Retriever-ST-late at revision 023be2a8ab9d797f5aa76f5bf8b5dde78d819659. Source data: Hugging Face dataset vidore/vidore_v3_computer_science at revision d5cc75883d92e294f0c0fc2662551c9708a06ebc, configs corpus / queries / qrels, split test, loaded with datasets. Document, query and qrel ids are the source's own ids… See the full description on the dataset page: https://huggingface.co/datasets/robro612/vidore3_computerscience_neomme_260m_li.

sourceHugging Facecc-by-4.0updated 14h agoView on Hugging Face
0likes
Dataset Card

vidore3computerscienceneomme260mli

Multi-vector (late-interaction) embeddings of ViDoRe computerscience (vidore/computerscience), encoded with [Hcompany/NeoMME-260M-Retriever-ST-late](https://huggingface.co/Hcompany/NeoMME-260M-Retriever-ST-late) at revision 023be2a8ab9d797f5aa76f5bf8b5dde78d819659.

Source data: Hugging Face dataset `vidore/vidore_v3_computer_science` at revision d5cc75883d92e294f0c0fc2662551c9708a06ebc, configs corpus / queries / qrels, split test, loaded with datasets. Document, query and qrel ids are the source's own ids, unchanged.

Every document is one variable-length set of 128-d vectors; every query is one variable-length set of 128-d vectors. Documents and queries are stored at different precisions (fp16 and fp32 respectively), see Encoding.

Files

filedtypeshapecontents
documents.npyfloat16 (<f2)[4,441,760, 128]every document vector, concatenated document by document (1,084.4 MiB)
doclens.npyint32[1,360]vectors per document; cumsum gives offsets
doc_ids.npy<U4[1,360]original document ids
queries.npyfloat32 (<f4)[1,290, 130, 128]query vectors, zero-padded at the end (81.9 MiB)
query_lens.npyint32[1,290]true vectors per query, before padding
queries_ids.npy<U4[1,290]original query ids
qrels.test.tsvtext6,294 rowsTREC qrels, qid \t 0 \t docid \t relevance, no header
gt_top100.tsvtext129,000 rowsexact MaxSim top-100, see below

All positional indices (gt_top100.tsv, and the row order of every .npy file) refer to the order of doc_ids.npy and queries_ids.npy. Reordering either file invalidates gt_top100.tsv.

Statistics

documents1,360
document vectors4,441,760
vectors per document (min / median / mean / max)3266 / 3266 / 3266.0 / 3266
queries1,290
vectors per query (min / median / max)15 / 32 / 130
queries with at least one qrel1,290
qrels rows6,294
embedding dimension128

Encoding

modelHcompany/NeoMME-260M-Retriever-ST-late
model revision023be2a8ab9d797f5aa76f5bf8b5dde78d819659
librarysentence-transformers 6.0.1 MultiVectorEncoder (transformers 5.17.0, torch 2.11.0+cu126)
document compute dtypefloat16 (model weights loaded at this dtype for the document pass)
document storage dtypefp16
query compute dtypefloat32 (model weights loaded at this dtype for the query pass)
query storage dtypefp32
normalizationL2, by the model's own Normalize module, before the storage cast
document truncationnone (document_length unset; model limit 16,384 tokens, longest document here 3,266 vectors, so nothing hit it)
query truncationnone (checkpoint default; query_length unset)
document skiplistnone (empty skiplist_words): every document token is kept, so doclens is the real token count
document inputpage image, one vector per image patch plus layout tokens; processor default resizing
query inputquery text, stripped of surrounding whitespace, formatted by the model's own query prompt/template
query vectorsevery vector the model emits for the query is kept, including any query-expansion tokens its template adds; query_lens counts them all
document paddingnone: documents.npy holds real vectors only, sum(doclens) == n_tokens
query paddingrows at or beyond query_lens[i] in queries.npy[i] are exactly zero
token_idsnot provided: image-patch vectors have no vocabulary ids (only placeholder ids)

Ground truth: gt_top100.tsv

Exact brute-force MaxSim top-100 per query over the full corpus, from the vectors in this repo. No header; tab-separated qidx docidx rank score:

  • —qidx: 0-based row into queries_ids.npy / queries.npy
  • —docidx: 0-based position into doc_ids.npy / doclens.npy
  • —rank: 1-based, descending score
  • —score: sum over the query's query_lens[qidx] vectors of max over the document's vectors of the dot product, computed in fp32 with the fp16 document vectors upcast to fp32. Expansion vectors are included in the sum. Printed to 6 decimals.

Self-matches are included. 92 of 1,290 queries are themselves corpus documents with the same id and retrieve that document (typically at rank 1). This is the raw nearest-neighbour list; BEIR's evaluation drops such pairs (ignore_identical_ids), so exclude them before scoring against qrels.

Retrieval quality

Sanity check of the vectors, not a leaderboard number: exact MaxSim over the full corpus scored against qrels.test.tsv with ir_measures, with query-id == doc-id pairs dropped as BEIR does.

nDCG@10Recall@100MRR@10MAP
0.67450.94040.81790.5996

Loading

python
import numpy as np

documents = np.load("documents.npy", mmap_mode="r")      # [n_tokens, 128] float16
doclens = np.load("doclens.npy")                         # [n_docs] int32
offsets = np.concatenate([[0], np.cumsum(doclens)])
doc_ids = np.load("doc_ids.npy")                         # [n_docs] str

def document(i):
    return documents[offsets[i]:offsets[i + 1]]           # [doclens[i], 128]

queries = np.load("queries.npy")                         # [n_queries, 130, 128] float32
query_lens = np.load("query_lens.npy")                   # [n_queries] int32
query_ids = np.load("queries_ids.npy")                   # [n_queries] str

def query(j):
    return queries[j, :query_lens[j]]                     # [query_lens[j], 128]

def maxsim(q, d):
    return (q @ d.astype(np.float32).T).max(axis=1).sum()

Validation

Checks run by the exporter on the files exactly as written here:

  • —✅ file set — missing=[] extra=[]
  • —✅ documents.npy dtype/shape — <f2 (4441760, 128)
  • —✅ doclens.npy dtype/shape — <i4 (1360,)
  • —✅ doc_ids.npy is a string array — <U4 (1360,)
  • —✅ queries.npy dtype/shape — <f4 (1290, 130, 128)
  • —✅ query_lens.npy dtype/shape — <i4 (1290,)
  • —✅ queries_ids.npy is a string array — <U4 (1290,)
  • —✅ sum(doclens) == n_tokens — 4441760 vs 4441760
  • —✅ no empty documents — min doclen 3266
  • —✅ len(doc_ids) == len(doclens) == corpus size — 1360, 1360, 1360
  • —✅ doc_ids unique and in corpus order
  • —✅ query arrays aligned — 1290, 1290, 1290
  • —✅ queries_ids in dataset order
  • —✅ max(query_lens) == queries.shape[1]
  • —✅ query padding is exactly zero
  • —✅ doc and query dim agree — 128 / 128
  • —✅ document vectors unit-norm (100k sample) — norm range [0.9995, 1.0006]
  • —✅ query vectors unit-norm — norm range [1.000000, 1.000000]
  • —✅ all vectors finite
  • —✅ qrels.test.tsv is 4-col TREC matching the dataset — 6294 rows
  • —✅ gt_top100.tsv has k rows per query — 129000 rows, k=100
  • —✅ gt rows grouped by qidx with ranks 1..k and descending scores
  • —✅ gt indices in range
  • —✅ gt scores reproduce from these files (16 queries, fp32 MaxSim) — max |diff| 1.14e-05
  • —✅ gt top-k is exact (no excluded doc outscores rank k) — max margin -4.13e-04
  • —✅ re-encoded docs match stored rows (4 docs) — lengths match, min token cosine 0.9950
  • —✅ re-encoded queries match stored rows (8 queries) — lengths match, min token cosine 1.000000

Provenance

exported2026-09-23
hardwareTesla V100S-PCIE-32GB