robro612/vidore3_energy_neomme_260m_li
vidore3_energy_neomme_260m_li Multi-vector (late-interaction) embeddings of ViDoRe energy (vidore/energy), encoded with Hcompany/NeoMME-260M-Retriever-ST-late at revision 023be2a8ab9d797f5aa76f5bf8b5dde78d819659. Source data: Hugging Face dataset vidore/vidore_v3_energy at revision caec06d3c73434d635f710f93bcd898331c59f20, configs corpus / queries / qrels, split test, loaded with datasets. Document, query and qrel ids are the source's own ids, unchanged. Every document is one… See the full description on the dataset page: https://huggingface.co/datasets/robro612/vidore3_energy_neomme_260m_li.
vidore3energyneomme260mli
Multi-vector (late-interaction) embeddings of ViDoRe energy (vidore/energy), encoded with [Hcompany/NeoMME-260M-Retriever-ST-late](https://huggingface.co/Hcompany/NeoMME-260M-Retriever-ST-late) at revision 023be2a8ab9d797f5aa76f5bf8b5dde78d819659.
Source data: Hugging Face dataset `vidore/vidore_v3_energy` at revision caec06d3c73434d635f710f93bcd898331c59f20, configs corpus / queries / qrels, split test, loaded with datasets. Document, query and qrel ids are the source's own ids, unchanged.
Every document is one variable-length set of 128-d vectors; every query is one variable-length set of 128-d vectors. Documents and queries are stored at different precisions (fp16 and fp32 respectively), see Encoding.
Files
All positional indices (gt_top100.tsv, and the row order of every .npy file) refer to the order of doc_ids.npy and queries_ids.npy. Reordering either file invalidates gt_top100.tsv.
Statistics
Encoding
Ground truth: gt_top100.tsv
Exact brute-force MaxSim top-100 per query over the full corpus, from the vectors in this repo. No header; tab-separated qidx docidx rank score:
qidx: 0-based row intoqueries_ids.npy/queries.npydocidx: 0-based position intodoc_ids.npy/doclens.npyrank: 1-based, descending scorescore:sum over the query's query_lens[qidx] vectors of max over the document's vectors of the dot product, computed in fp32 with the fp16 document vectors upcast to fp32. Expansion vectors are included in the sum. Printed to 6 decimals.
Self-matches are included. 71 of 1,848 queries are themselves corpus documents with the same id and retrieve that document (typically at rank 1). This is the raw nearest-neighbour list; BEIR's evaluation drops such pairs (ignore_identical_ids), so exclude them before scoring against qrels.
Retrieval quality
Sanity check of the vectors, not a leaderboard number: exact MaxSim over the full corpus scored against qrels.test.tsv with ir_measures, with query-id == doc-id pairs dropped as BEIR does.
Loading
import numpy as np
documents = np.load("documents.npy", mmap_mode="r") # [n_tokens, 128] float16
doclens = np.load("doclens.npy") # [n_docs] int32
offsets = np.concatenate([[0], np.cumsum(doclens)])
doc_ids = np.load("doc_ids.npy") # [n_docs] str
def document(i):
return documents[offsets[i]:offsets[i + 1]] # [doclens[i], 128]
queries = np.load("queries.npy") # [n_queries, 70, 128] float32
query_lens = np.load("query_lens.npy") # [n_queries] int32
query_ids = np.load("queries_ids.npy") # [n_queries] str
def query(j):
return queries[j, :query_lens[j]] # [query_lens[j], 128]
def maxsim(q, d):
return (q @ d.astype(np.float32).T).max(axis=1).sum()Validation
Checks run by the exporter on the files exactly as written here:
- ✅ file set — missing=[] extra=[]
- ✅ documents.npy dtype/shape — <f2 (6466875, 128)
- ✅ doclens.npy dtype/shape — <i4 (2225,)
- ✅ doc_ids.npy is a string array — <U4 (2225,)
- ✅ queries.npy dtype/shape — <f4 (1848, 70, 128)
- ✅ query_lens.npy dtype/shape — <i4 (1848,)
- ✅ queries_ids.npy is a string array — <U4 (1848,)
- ✅ sum(doclens) == n_tokens — 6466875 vs 6466875
- ✅ no empty documents — min doclen 1001
- ✅ len(doc_ids) == len(doclens) == corpus size — 2225, 2225, 2225
- ✅ doc_ids unique and in corpus order
- ✅ query arrays aligned — 1848, 1848, 1848
- ✅ queries_ids in dataset order
- ✅ max(query_lens) == queries.shape[1]
- ✅ query padding is exactly zero
- ✅ doc and query dim agree — 128 / 128
- ✅ document vectors unit-norm (100k sample) — norm range [0.9994, 1.0006]
- ✅ query vectors unit-norm — norm range [1.000000, 1.000000]
- ✅ all vectors finite
- ✅ qrels.test.tsv is 4-col TREC matching the dataset — 6618 rows
- ✅ gt_top100.tsv has k rows per query — 184800 rows, k=100
- ✅ gt rows grouped by qidx with ranks 1..k and descending scores
- ✅ gt indices in range
- ✅ gt scores reproduce from these files (16 queries, fp32 MaxSim) — max |diff| 1.14e-05
- ✅ gt top-k is exact (no excluded doc outscores rank k) — max margin -3.83e-04
- ✅ re-encoded docs match stored rows (4 docs) — lengths match, min token cosine 0.9990
- ✅ re-encoded queries match stored rows (8 queries) — lengths match, min token cosine 1.000000
