CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /trec-covid TRECCOVID An MTEB dataset Massive Text Embedding Benchmark TRECCOVID is an ad-hoc search challenge based on the COVID-19 dataset containing scientific articles related to the COVID-19 pandemic. Task category t2t Domains Medical, Academic, Written Reference https://ir.nist.gov/covidSubmit/index.html How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["TRECCOVID"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/trec-covid.texttext-retrieval100K<n<1M5 likes7.9k downloads7mo agoHugging Face02SetFit /TREC-QC TREC Question Classification Question classification in coarse and fine-grained categories. Source: Experimental Data for Question Classification Xin Li, Dan Roth, Learning Question Classifiers. COLING'02, Aug., 2002. tabular1K<n<10K0 likes3.2k downloads5y agoHugging Face03TREC-AToMiC /AToMiC-Baselines AToMiC Prebuilt Indexes Example Usage: Reproduction Toolkits: https://github.com/TREC-AToMiC/AToMiC/tree/main/examples/dense_retriever_baselines # Skip the encode and index steps, search with the prebuilt indexes and topics directly python search.py \ --topics topics/openai.clip-vit-base-patch32.text.validation \ --index indexes/openai.clip-vit-base-patch32.image.faiss.flat \ --hits 1000 \ --output… See the full description on the dataset page: https://huggingface.co/datasets/TREC-AToMiC/AToMiC-Baselines.textn<1K1 likes659 downloads3y agoHugging Face04trec-ragtime /ragtime1 RAGTIME1 Collection This dataset contains the documents for TREC RAGTIME Track. Please refer to the website for the details of the task. RAGTIME is a multilingual RAG task, which expects the participating system to retrieve relevant documents from all four languages and synthesize a response with citation to the report request. For convenience, we separate the documents by their languages into four .jsonl files. However, they are intended to be used as a whole set. The documents… See the full description on the dataset page: https://huggingface.co/datasets/trec-ragtime/ragtime1.texttext-retrieval1M<n<10M0 likes232 downloads10mo agoHugging Face05hyesunyun /liveqa_medical_trec2017 Dataset Card for LiveQA Medical from TREC 2017 The LiveQA'17 medical task focuses on consumer health question answering. Consumer health questions were received by the U.S. National Library of Medicine (NLM). The dataset consists of constructed medical question-answer pairs for training and testing, with additional annotations that can be used to develop question analysis and question answering systems. Please refer to our overview paper for more information about the constructed… See the full description on the dataset page: https://huggingface.co/datasets/hyesunyun/liveqa_medical_trec2017.textquestion-answeringn<1K8 likes205 downloads3y agoHugging Face06michaeldinzinger /merged-trecdltexttext-retrieval1M<n<10M0 likes153 downloads1y agoHugging Face07whybe-choi /trec-dl-2019 TRECDL2019 An MTEB dataset Massive Text Embedding Benchmark TREC Deep Learning Track 2019 passage ranking task. The task involves retrieving relevant passages from the MS MARCO collection given web search queries. Queries have multi-graded relevance judgments. Task categoryt2t Domains Encyclopaedic, Academic, Blog, News, Medical, Government, Reviews, Non-fiction, Social, Web Reference https://microsoft.github.io/msmarco/TREC-Deep-Learning-2019 How to… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/trec-dl-2019.texttext-retrieval1M<n<10M1 likes149 downloads11mo agoHugging Face08liuqi6777 /trec-newstexttext-retrieval100K<n<1M0 likes140 downloads1y agoHugging Face09TREC-AToMiC /atomic2024image100K<n<1M0 likes127 downloads2y agoHugging Face10harisarang /benchmark-trec-covidtext100K<n<1M0 likes86 downloads10mo agoHugging Face11whybe-choi /trec-dl-2020 TRECDL2020 An MTEB dataset Massive Text Embedding Benchmark TREC Deep Learning Track 2020 passage ranking task. The task involves retrieving relevant passages from the MS MARCO collection given web search queries. Queries have multi-graded relevance judgments. Task categoryt2t Domains Encyclopaedic, Academic, Blog, News, Medical, Government, Reviews, Non-fiction, Social, Web Reference https://microsoft.github.io/msmarco/TREC-Deep-Learning-2020 How to… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/trec-dl-2020.texttext-retrieval1M<n<10M1 likes74 downloads11mo agoHugging Face12BeIR /trec-news-generated-queries Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/trec-news-generated-queries.texttext-retrieval1M<n<10M4 likes67 downloads4y agoHugging Face13BeIR /trec-covid-generated-queries Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/trec-covid-generated-queries.texttext-retrieval100K<n<1M1 likes63 downloads4y agoHugging Face14MCINext /trec-covid-fa-v2text100K<n<1M0 likes50 downloads1y agoHugging Face15TREC-AToMiC /atomic2023-small_text2imageimage10K<n<100K1 likes40 downloads2y agoHugging Face16UTokyo-Yokoya-Lab /trec-covid-CSR-L TRECCOVID-CodeSwitching An MTEB dataset Massive Text Embedding Benchmark Code-switching version of mteb/trec-covid, with queries rewritten in Chinese-English and Japanese-English code-switching styles. Dataset Structure The dataset contains the following configurations: From original dataset (unchanged): corpus: Original corpus documents default: Original relevance judgments (qrels) Code-switching additions: queries_zh_en: Chinese-English code-switching queries… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/trec-covid-CSR-L.texttext-retrieval100K<n<1M0 likes38 downloads6mo agoHugging Face17nreimers /trec-covid-generated-queriestext10K<n<100K0 likes36 downloads5y agoHugging Face18kaengreg /rus-trec-covidtext100K<n<1M0 likes30 downloads2y agoHugging Face19nreimers /trec-covidThis is the corpus file from the BEIR benchmark for the TREC-COVID 19 dataset. text100K<n<1M1 likes28 downloads5y agoHugging Face20trec-ragtime /ragtime2 RAGTIME2 Collection This dataset contains the documents for TREC RAGTIME Track 2026. Please refer to the website for the details of the task. RAGTIME is a multilingual RAG task, which expects the participating system to retrieve relevant documents from all four languages and synthesize a response with citation to the report request. For convenience, we separate the documents by their languages into four .jsonl files. However, they are intended to be used as a whole set. The… See the full description on the dataset page: https://huggingface.co/datasets/trec-ragtime/ragtime2.texttext-retrieval1M<n<10M0 likes28 downloads5mo agoHugging Face21nthakur /gpl-trec-covidtext100K<n<1M1 likes27 downloads3y agoHugging Face22irds /trec_cast_offsets Dataset Card for Dataset Name This is a complement to the TREC CaST (2020-22) datasets, with pre-computed offset relative to the original files. text10M<n<100M0 likes27 downloads3y agoHugging Face23liuqi6777 /trec_dl19texttext-retrieval1M<n<10M0 likes27 downloads1y agoHugging Face24Leooyii /manyshots_trectext1K<n<10K0 likes25 downloads2y agoHugging Face25clips /beir-nl-trec-covid Dataset Card for BEIR-NL Benchmark Dataset Summary BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB). BEIR-NL contains the following tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-trec-covid.texttext-retrieval100K<n<1M0 likes25 downloads2y agoHugging Face26Nithish2410 /benchmark-trec-covidtext100K<n<1M0 likes24 downloads7mo agoHugging Face27DomLoyer /trec-ap88-90-corpus TREC AP88-90 Corpus Dataset Summary Le TREC AP88-90 Corpus est un corpus de documents Associated Press couvrant la periode 1988-1990. Il est destine a des experiments en recherche d'information, en ranking, et en evaluation de systemes de retrieval. Dataset Description Overview Ce depot contient une version preparee du corpus TREC AP88-90 pour des usages de recherche et d'experimentation.Le contenu est… See the full description on the dataset page: https://huggingface.co/datasets/DomLoyer/trec-ap88-90-corpus.text100K<n<1M1 likes22 downloads6mo agoHugging Face28MCINext /trec-covid-fa Dataset Summary TRECCOVID-Fa is a Persian (Farsi) dataset designed for the Retrieval task, specifically focusing on ad-hoc search for COVID-19-related scientific information. It is a translated version of the English dataset from the TREC-COVID shared task, included in the BEIR benchmark, and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) under the BEIR-Fa collection. Language(s): Persian (Farsi) Task(s): Retrieval (Ad-hoc Search, COVID-19 Information Retrieval)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/trec-covid-fa.text100K<n<1M0 likes19 downloads1y agoHugging Face29liuqi6777 /trec_dl20texttext-retrieval1M<n<10M0 likes17 downloads1y agoHugging Face30jordane95 /trec-dl-2019-querytextn<1K0 likes11 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.