CoolFace
Datasetpublic

tuskanny/nfcorpus_lateon

NFCorpus, LateOn Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with LateOn, in the TACHIOM multivector format. Source BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding 3,633 documents, 323 queries, 12,334 qrels Text given to the encoder for each document: title + " " + text (BEIR title and body joined by a… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_lateon.

sourceHugging Faceupdated 4d agoView on Hugging Face
0likes45downloads
filedoc_ids.npy114 KBdownload
filedoclens.npy14 KBdownload
filedocuments.npy210.6 MBdownload
filequeries_ids.npy13 KBdownload
filequeries.npy3.9 MBdownload
filequery_lens.npy1 KBdownload
filetoken_ids.npy3.3 MBdownload

tuskanny/nfcorpus_lateon · main · files are served by the source, never re-hosted here