CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BeIR /msmarco-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/msmarco-qrels.tabulartext-retrieval100K<n<1M1 likes1.1k downloads4y agoHugging Face02bdjafer /msmarco-yesnotabular10K<n<100K0 likes279 downloads4y agoHugging Face03tuskanny /ms_marco_colbertv2 MS MARCO v1 Passage, ColBERTv2 Token-level (late-interaction) ColBERTv2 embeddings of the MS MARCO v1 passage collection and the dev/small queries. Source Collection: MS MARCO v1 passage (ir_datasets msmarco-passage), 8,841,823 passages Queries: dev/small, 6,980 queries and 7,437 qrels (msmarco-passage/dev/small) Document order: passage id order (row i is pid i) Encoding Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/ms_marco_colbertv2.tabulartext-retrieval1K<n<10K0 likes158 downloads2h agoHugging Face04AnonymousUser2026 /ms_marco_cocondensertabular1K<n<10K0 likes46 downloads11mo agoHugging Face05tuskanny /kannolo-msmarco-spladetabular1K<n<10K0 likes28 downloads1y agoHugging Face06clarin-knext /msmarco-pl-qrelsPart of BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language. Link to arxiv: https://arxiv.org/pdf/2305.19840.pdf Contact: konrad.wojtasik@pwr.edu.pl tabular100K<n<1M0 likes18 downloads3y agoHugging Face07freethenation /msmarco-sub MS MARCO Subset Purpose This subset provides a standardized benchmark for evaluating sparse model performance on MS MARCO data, with negative examples pooled using BM25 retrieval. Creation This subset was created using the make_beir_subset.py script with the following command: python ./make_beir_subset.py --dataset msmarco --es-host http://localhost:9200 --force-reindex --split dev Parameters Used Dataset: msmarco Split: dev (development set) ES… See the full description on the dataset page: https://huggingface.co/datasets/freethenation/msmarco-sub.tabular1K<n<10K0 likes17 downloads11mo agoHugging Face08AnonymousUser2026 /ms_marco_inferencelesstabular1K<n<10K0 likes16 downloads11mo agoHugging Face09unicamp-dl /InRanker-msmarcotabular10M<n<100M0 likes10 downloads3y agoHugging Face10saracandu /msmarco_modifiedtabular10K<n<100K0 likes9 downloads2y agoHugging Face11tuskanny /kannolo-msmarco-cocondensertabular10K<n<100K0 likes5 downloads4mo agoHugging Face12shaneHF /CIIR_MSMARCOtabular10M<n<100M0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.