CoolFace
Datasetpublic

tuskanny/nfcorpus_colbertv2

NFCorpus, ColBERTv2 Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format. Source BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding 3,633 documents, 323 queries, 12,334 qrels Text given to the encoder for each document: title + " " + text (BEIR title and body joined by a… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_colbertv2.

sourceHugging Faceupdated 1d agoView on Hugging Face
0likes39downloads
filedoc_ids.npy114 KBdownload
filedoclens.npy14 KBdownload
filedocuments.npy137.0 MBdownload
filequeries_ids.npy13 KBdownload
filequeries.npy5.0 MBdownload
filequery_lens.npy1 KBdownload
filetoken_ids.npy2.1 MBdownload

tuskanny/nfcorpus_colbertv2 · main · files are served by the source, never re-hosted here