CoolFace
Datasetpublic

tuskanny/fiqa_lateon

FiQA-2018, LateOn Token-level (late-interaction) embeddings of the BEIR FiQA-2018 corpus and queries, encoded with LateOn, in the TACHIOM multivector format. Source BEIR FiQA-2018, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/fiqa/test); PyLate only did the encoding 57,638 documents, 648 queries, 1,706 qrels Text given to the encoder for each document: the passage text (FiQA documents have no title). The text… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/fiqa_lateon.

sourceHugging Faceupdated 3d agoView on Hugging Face
0likes38downloads
filedoc_ids.npy1.3 MBdownload
filedoclens.npy225 KBdownload
filedocuments.npy1.83 GBdownload
filequeries_ids.npy13 KBdownload
filequeries.npy10.1 MBdownload
filequery_lens.npy3 KBdownload
filetoken_ids.npy29.4 MBdownload

tuskanny/fiqa_lateon · main · files are served by the source, never re-hosted here