datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trec-covid
TRECCOVID
An MTEB dataset
Massive Text Embedding Benchmark
TRECCOVID is an ad-hoc search challenge based on the COVID-19 dataset containing scientific articles related to the COVID-19 pandemic.
Task category
t2t
Domains
Medical, Academic, Written
Reference
https://ir.nist.gov/covidSubmit/index.html
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TRECCOVID"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/trec-covid.trec-covid
Dataset Card for BEIR Benchmark
trec-covid is one of the datasets from the Bio-Medical Retrieval task within BEIR, measuring scientific article retrieval for a given query on COVID-19.
Dataset Summary
BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks.
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/trec-covid.trec-covid-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/trec-covid-qrels.trec-covid-decontaminated
trec-covid (Decontaminated)
A decontaminated version of the trec-covid dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed.
Decontamination methodology
Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files):
Pass 1: Exact hash matching
All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed with… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/trec-covid-decontaminated.TRECCOVID-PL
TRECCOVID-PL
An MTEB dataset
Massive Text Embedding Benchmark
TRECCOVID is an ad-hoc search challenge based on the COVID-19 dataset containing scientific articles related to the COVID-19 pandemic.
Task category
t2t
Domains
Academic, Medical, Non-fiction, Written
Reference
https://ir.nist.gov/covidSubmit/index.html
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/TRECCOVID-PL.benchmark-trec-covidtrec-covid-vn
TRECCOVID-VN
An MTEB dataset
Massive Text Embedding Benchmark
A translated dataset from TRECCOVID is an ad-hoc search challenge based on the COVID-19 dataset containing scientific articles related to the COVID-19 pandemic. The process of creating the VN-MTEB (Vietnamese Massive Text Embedding Benchmark) from English samples involves a new automated system: - The system uses large language models (LLMs), specifically Coherence's Aya model, for translation. - Applies advanced embedding… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/trec-covid-vn.beir-trec-covid
TRECCOVID — BEIR, unified schema
A normalised copy of the dataset behind the mteb task TRECCOVID, one of the tasks of the BEIR benchmark as mteb defines it. Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/trec-covid @ bb9466bac815 (the revision pinned in mteb)
Domain · languages
biomedical · eng
Queries / documents / qrels
50 / 171,332 / 66,336… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-trec-covid.trec-covid_CS-MTEB
TREC-COVID CS-MTEB
An MTEB dataset
Massive Text Embedding Benchmark
Code-switching version of mteb/trec-covid, with queries rewritten in Chinese-English, Japanese-English, German-English, Spanish-English, Korean-English, French-English, Italian-English, Portuguese-English, Dutch-English code-switching styles.
Dataset Structure
The dataset contains the following configurations:
From original dataset (unchanged):
corpus: Original corpus documents
default: Original relevance… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/trec-covid_CS-MTEB.trec_covid_toyset_pairtrec-covid-generated-queries
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/trec-covid-generated-queries.trec-covid-fa-v2trec-covid-decontaminated
trec-covid-decontaminated (MTEB layout)
Repackaging of lightonai/trec-covid-decontaminated
into the layout expected by MTEB:
a default config holding the qrels (one split per evaluation split), alongside
corpus and queries configs.
Rows and columns are copied verbatim from the source dataset; only the
config/split packaging differs. All credit for the decontaminated data belongs to
LightOn AI, and to the original BEIR
authors for the underlying benchmark.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/trec-covid-decontaminated.trec-covid-CSR-L
TRECCOVID-CodeSwitching
An MTEB dataset
Massive Text Embedding Benchmark
Code-switching version of mteb/trec-covid, with queries rewritten in Chinese-English and Japanese-English code-switching styles.
Dataset Structure
The dataset contains the following configurations:
From original dataset (unchanged):
corpus: Original corpus documents
default: Original relevance judgments (qrels)
Code-switching additions:
queries_zh_en: Chinese-English code-switching queries… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/trec-covid-CSR-L.trec-covid-generated-queriestrec-covid__openai_ada2rus-trec-covidTRECCOVID-Fa
TRECCOVID-Fa
An MTEB dataset
Massive Text Embedding Benchmark
TRECCOVID-Fa
Task category
t2t
Domains
Medical
Reference
https://huggingface.co/datasets/MCINext/trec-covid-fa
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("TRECCOVID-Fa")
evaluator = mteb.MTEB([task])
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on… See the full description on the dataset page: https://huggingface.co/datasets/mteb/TRECCOVID-Fa.trec-covidThis is the corpus file from the BEIR benchmark for the TREC-COVID 19 dataset.
gpl-trec-covidbeir-nl-trec-covid
Dataset Card for BEIR-NL Benchmark
Dataset Summary
BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB).
BEIR-NL contains the following tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-trec-covid.trec-covid-pl-qrelsPart of BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language.
Link to arxiv: https://arxiv.org/pdf/2305.19840.pdf
Contact: konrad.wojtasik@pwr.edu.pl
benchmark-trec-covidtrec-covid
Dataset Card for BEIR Benchmark
trec-covid is one of the datasets from the Bio-Medical Retrieval task within BEIR, measuring scientific article retrieval for a given query on COVID-19.
Dataset Summary
BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks.
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval:… See the full description on the dataset page: https://huggingface.co/datasets/chenxiaobin/trec-covid.rus-trec-covid-qrelstrec-covid-fa
Dataset Summary
TRECCOVID-Fa is a Persian (Farsi) dataset designed for the Retrieval task, specifically focusing on ad-hoc search for COVID-19-related scientific information. It is a translated version of the English dataset from the TREC-COVID shared task, included in the BEIR benchmark, and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) under the BEIR-Fa collection.
Language(s): Persian (Farsi)
Task(s): Retrieval (Ad-hoc Search, COVID-19 Information Retrieval)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/trec-covid-fa.treccovid_az-corpusTRECCOVID-NL
TRECCOVID-NL
An MTEB dataset
Massive Text Embedding Benchmark
TRECCOVID is an ad-hoc search challenge based on the COVID-19 dataset containing scientific articles related to the COVID-19 pandemic. TRECCOVID-NL is a Dutch translation.
Task category
t2t
Domains
Medical, Academic, Written
Reference
https://colab.research.google.com/drive/1R99rjeAGt8S9IfAIRR3wS052sNu3Bjo-#scrollTo=4HduGW6xHnrZ
How to evaluate on this task
You can evaluate an embedding… See the full description on the dataset page: https://huggingface.co/datasets/mteb/TRECCOVID-NL.trec_covidtreccovidThis dataset is a reranking-formatted version of the TREC-COVID dataset from the BEIR benchmark.
Original Dataset: TREC-COVID from NISTBEIR Version: BeIR/trec-covidLicense: Available for research purposesTask: COVID-19 scientific literature retrieval
