CoolFace
Datasetpublic

tuskanny/nfcorpus_modern_colbert

NFCorpus, GTE-ModernColBERT Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format. Source BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding 3,633 documents, 323 queries, 12,334 qrels Text given to the encoder for each document: title + " " + text (BEIR title and… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_modern_colbert.

sourceHugging Faceupdated 3d agoView on Hugging Face
0likes39downloads
filedoc_ids.npy114 KBdownload
filedoclens.npy14 KBdownload
filedocuments.npy210.6 MBdownload
filequeries_ids.npy13 KBdownload
filequeries.npy3.9 MBdownload
filequery_lens.npy1 KBdownload
filetoken_ids.npy3.3 MBdownload

tuskanny/nfcorpus_modern_colbert · main · files are served by the source, never re-hosted here