CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BeIR /scifact Dataset Card for BEIR Benchmark scifact is one of the datasets from the Fact Checking task within BEIR, measuring scientific article retrieval for a given scientific claim. Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/scifact.textzero-shot-classification1K<n<10K9 likes15k downloads6mo agoHugging Face02BeIR /scidocs Dataset Card for BEIR Benchmark scidocs is one of the datasets from the Citation Prediction task within BEIR, measuring cited scientific article retrieval for a given scientific title. Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/scidocs.textzero-shot-classification10K<n<100K10 likes12k downloads6mo agoHugging Face03BeIR /scifact-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/scifact-qrels.tabulartext-retrieval1K<n<10K1 likes11k downloads4y agoHugging Face04pinecone /msmarco-beir-e50 likes7.5k downloads1y agoHugging Face05BeIR /nfcorpus Dataset Card for BEIR Benchmark nfcorpus is one of the datasets from the Bio-Medical Retrieval task within BEIR, measuring the retrieval of scientific articles for a given query about a nutritional fact. Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/nfcorpus.textzero-shot-classification1K<n<10K7 likes6.2k downloads6mo agoHugging Face06clips /beir-nl-cqadupstack Dataset Card for BEIR-NL Benchmark Dataset Summary BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB). BEIR-NL contains the following tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-cqadupstack.texttext-retrieval100K<n<1M0 likes5.3k downloads2y agoHugging Face07BeIR /nfcorpus-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/nfcorpus-qrels.texttext-retrieval100K<n<1M0 likes5.2k downloads4y agoHugging Face08BeIR /fiqa Dataset Card for BEIR Benchmark fiqa is one of the datasets from the Question Answering task within BEIR, measuring financial article retrieval for a given financial query. Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/fiqa.textzero-shot-classification10K<n<100K16 likes4.1k downloads6mo agoHugging Face09castorini /prebuilt-indexes-beir Prebuilt Indexes for BEIR Available indexes: Lucene Flat beir-v1.0.0-trec-covid.bge-base-en-v1.5.flat [readme] Lucene flat index of BEIR collection 'trec-covid' encoded by BGE-base-en-v1.5. beir-v1.0.0-bioasq.bge-base-en-v1.5.flat [readme] Lucene flat index of BEIR collection 'bioasq' encoded by BGE-base-en-v1.5. beir-v1.0.0-nfcorpus.bge-base-en-v1.5.flat [readme] Lucene flat index of BEIR collection 'nfcorpus' encoded by BGE-base-en-v1.5. beir-v1.0.0-nq.bge-base-en-v1.5.flat… See the full description on the dataset page: https://huggingface.co/datasets/castorini/prebuilt-indexes-beir.1 likes3.9k downloads1y agoHugging Face10BeIR /fiqa-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/fiqa-qrels.tabulartext-retrieval10K<n<100K0 likes3.4k downloads4y agoHugging Face11TIGER-Lab /M-BEIR UniIR: Training and Benchmarking Universal Multimodal Information Retrievers (ECCV 2024) 🌐 Homepage | 🤗 Model(UniIR Checkpoints) | 🤗 Paper | 📖 arXiv | GitHub How to download the M-BEIR Dataset 🔔News 🔥[2023-12-21]: Our M-BEIR Benchmark is now available for use. Dataset Summary M-BEIR, the Multimodal BEnchmark for Instructed Retrieval, is a comprehensive large-scale retrieval benchmark designed to train and evaluate unified multimodal retrieval… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/M-BEIR.texttext-retrieval1M<n<10M27 likes2.4k downloads2y agoHugging Face12pinecone /msmarco-beir-constbert0 likes2.4k downloads1y agoHugging Face13BeIR /msmarco Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. This msmarco subset is part of BEIR. Languages All tasks are in English (en). Dataset Structure This dataset uses the standard BEIR retrieval layout and includes: corpus: one row per document with _id, title, text queries: one row per query with _id, title, text Data Fields _id… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/msmarco.textzero-shot-classification1M<n<10M14 likes2.3k downloads6mo agoHugging Face14CohereLabs /beir-embed-english-v3 BEIR embeddings with Cohere embed-english-v3.0 model This datasets contains all query & document embeddings for BEIR, embedded with the Cohere embed-english-v3.0 embedding model. Overview of datasets This repository hosts all 18 datasets from BEIR, including query and document embeddings. The following table gives an overview of the available datasets. See the next section how to load the individual datasets. Dataset nDCG@10 #Documents arguana 53.98 8,674… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/beir-embed-english-v3.text10M<n<100M9 likes2.1k downloads6mo agoHugging Face15BeIR /trec-covid Dataset Card for BEIR Benchmark trec-covid is one of the datasets from the Bio-Medical Retrieval task within BEIR, measuring scientific article retrieval for a given query on COVID-19. Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/trec-covid.textzero-shot-classification100K<n<1M7 likes2.1k downloads6mo agoHugging Face16vidore /arxivqa_test_subsampled_beirBEIR version of vidore/arxivqa_test_subsampled. imagedocument-question-answering1K<n<10K1 likes1.9k downloads1y agoHugging Face17vidore /docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.8k downloads1y agoHugging Face18vidore /tabfquad_test_subsampled_beirBEIR version of vidore/tabfquad_test_subsampled. imagedocument-question-answeringn<1K0 likes1.7k downloads1y agoHugging Face19vidore /infovqa_test_subsampled_beirBEIR version of vidore/infovqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.7k downloads1y agoHugging Face20vidore /tatdqa_test_beirBEIR version of vidore/tatdqa_test. imagedocument-question-answering1K<n<10K0 likes1.7k downloads1y agoHugging Face21vidore /shiftproject_test_beirBEIR version of vidore/shiftproject_test. imagedocument-question-answering1K<n<10K0 likes1.7k downloads1y agoHugging Face22vidore /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes1.6k downloads1y agoHugging Face23vidore /syntheticDocQA_energy_test_beirBEIR version of vidore/syntheticDocQA_energy_test. imagedocument-question-answering1K<n<10K0 likes1.6k downloads1y agoHugging Face24vidore /syntheticDocQA_healthcare_industry_test_beirBEIR version of vidore/syntheticDocQA_healthcare_industry_test. imagedocument-question-answering1K<n<10K0 likes1.6k downloads1y agoHugging Face25BeIR /arguana Dataset Card for BEIR Benchmark arguana is one of the datasets from the Argument Retrieval task within BEIR, evaluating retrieving counterarguments for a given argument as query. NOTE: ArguAna has queries also incorporated within the corpus, so you should remove the same query_id if present within corpus during inference (implemented in BEIR) Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks.… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/arguana.textzero-shot-classification10K<n<100K4 likes1.6k downloads6mo agoHugging Face26vidore /syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test. imagedocument-question-answering1K<n<10K1 likes1.6k downloads1y agoHugging Face27BeIR /nq Dataset Card for BEIR Benchmark nq is one of the datasets from the Question Answering task within BEIR, measuring Wikipedia article retrieval for a web search query. Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/nq.textzero-shot-classification1M<n<10M4 likes1.6k downloads6mo agoHugging Face28BeIR /trec-covid-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/trec-covid-qrels.tabulartext-retrieval10K<n<100K1 likes1.5k downloads4y agoHugging Face29BeIR /quora Dataset Card for BEIR Benchmark quora is one of the datasets from the Duplicate Question Retrieval task within BEIR, measuring duplicate query retrieval for a given query. NOTE: ArguAna has queries also incorporated within the corpus, so you should remove the same query_id if present within the corpus during inference (implemented in BEIR) Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks.… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/quora.textzero-shot-classification100K<n<1M5 likes1.4k downloads6mo agoHugging Face30BeIR /arguana-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/arguana-qrels.texttext-retrieval1K<n<10K0 likes1.3k downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.