datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trec-covid
TRECCOVID
An MTEB dataset
Massive Text Embedding Benchmark
TRECCOVID is an ad-hoc search challenge based on the COVID-19 dataset containing scientific articles related to the COVID-19 pandemic.
Task category
t2t
Domains
Medical, Academic, Written
Reference
https://ir.nist.gov/covidSubmit/index.html
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TRECCOVID"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/trec-covid.covid19_emergency_event
Dataset Card for EXCEPTIUS Corpus
Dataset Summary
This dataset presents a new corpus of legislative documents from 8 European countries (Beglium, France, Hunary, Italy, Netherlands, Norway, Poland, UK) in 7 languages (Dutch, English, French, Hungarian, Italian, Norwegian Bokmål, Polish) manually annotated for exceptional measures against COVID-19. The annotation was done on the sentence level.
Supported Tasks and Leaderboards
The dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/covid19_emergency_event.ClusTREC-Covid
CLUSTREC-COVID: A Topical Clustering Benchmark for COVID-19 Scientific Research
Dataset Summary
CLUSTREC-COVID is a modified version of the TREC-COVID dataset, transformed into a topical clustering benchmark. The dataset consists of titles and abstracts from scientific papers about COVID-19 research, covering a diverse range of research topics. Each document in the dataset is assigned to a specific subtopic, making it ideal for use in document clustering and topic… See the full description on the dataset page: https://huggingface.co/datasets/Uri-ka/ClusTREC-Covid.covid19-ct-seg
COVID-19 CT Segmentation Dataset
Dataset Description
The COVID-19 CT Segmentation dataset for lung and COVID-19 infection segmentation from CT scans. This dataset contains CT scans with dense segmentation annotations.
Dataset Details
Modality: CT
Target: left lung, right lung, COVID-19 infection
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz",
"mask":… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/covid19-ct-seg.benchmark-trec-covidtrec-covid-generated-queries
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/trec-covid-generated-queries.trec-covid-fa-v2trec-covid-generated-queriestrec-covid-CSR-L
TRECCOVID-CodeSwitching
An MTEB dataset
Massive Text Embedding Benchmark
Code-switching version of mteb/trec-covid, with queries rewritten in Chinese-English and Japanese-English code-switching styles.
Dataset Structure
The dataset contains the following configurations:
From original dataset (unchanged):
corpus: Original corpus documents
default: Original relevance judgments (qrels)
Code-switching additions:
queries_zh_en: Chinese-English code-switching queries… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/trec-covid-CSR-L.COVID-QA-el-small
Dataset Card for COVID-QA-el-small
Dataset Summary
The COVID-QA-el-small dataset is a Greek-language subset of 826 examples derived from the COVID-QA-el dataset, translated using machine translation. The dataset follows the SQuADv1.1 fashion style.
The original dataset, COVID-QA: A Question Answering Dataset for COVID-19 (ACL 2020) contains 2,019 question-answer pairs annotated by volunteer biomedical experts on scientific literature about COVID-19.
Data… See the full description on the dataset page: https://huggingface.co/datasets/panosgriz/COVID-QA-el-small.trec-covidThis is the corpus file from the BEIR benchmark for the TREC-COVID 19 dataset.
gpl-trec-covidrus-trec-covidbeir-nl-trec-covid
Dataset Card for BEIR-NL Benchmark
Dataset Summary
BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB).
BEIR-NL contains the following tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-trec-covid.benchmark-trec-covidcovid_colbert_pgtr_golden
COVID ColBERT PGTR Golden
COVID queries and corpus from Nithish2410/covid_colbert_pgtr_golden, with the previous targets ignored and replaced by Qwen-reranked top-100 targets.
Contents
train.jsonl: 10,000 queries with 100 Qwen-reranked targets each.
items.jsonl: 171,332 COVID corpus passages.
Rerank Setup
Query source: existing query texts from the dataset.
Corpus source: existing items split from the dataset.
Candidate source: e5-base-v2… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/covid_colbert_pgtr_golden.covid-stpp
COVID-19 NJ STPP Benchmark Dataset
A benchmark-ready Spatio-Temporal Point Process (STPP) dataset derived from
COVID-19 case data for New Jersey, following the official split semantics of the
Neural STPP paper.
Dataset Description
Each record represents a sequence of COVID-19 case-report events for a
specific reporting unit on a specific date in 2020 (March–July).
Events capture the time and location of reported cases.
Source Format
Raw data was… See the full description on the dataset page: https://huggingface.co/datasets/seahorse-stpp/covid-stpp.trec-covid-fa
Dataset Summary
TRECCOVID-Fa is a Persian (Farsi) dataset designed for the Retrieval task, specifically focusing on ad-hoc search for COVID-19-related scientific information. It is a translated version of the English dataset from the TREC-COVID shared task, included in the BEIR benchmark, and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) under the BEIR-Fa collection.
Language(s): Persian (Farsi)
Task(s): Retrieval (Ad-hoc Search, COVID-19 Information Retrieval)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/trec-covid-fa.covid19-ct-seg
COVID-19 CT Segmentation Dataset
Dataset Description
The COVID-19 CT Segmentation dataset for lung and COVID-19 infection segmentation from CT scans. This dataset contains CT scans with dense segmentation annotations.
Dataset Details
Modality: CT
Target: left lung, right lung, COVID-19 infection
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz",
"mask":… See the full description on the dataset page: https://huggingface.co/datasets/Zhao-zi/covid19-ct-seg.lynx-70b-instruct-covidqa-generationscovidsample test dataset
nlu-covidFrench benchmark of NLU services for employee support use case during covid-19 pandemic.
These datasets were created by the Wikit team in order to compare the performances of NLU tools on the French language.
The dataset use case is employee support during the covid 19 pandemic. The intents were defined to answer department employees' questions on the evolution of work conditions related to the crisis.
The training_dataset.csv file contains training utterances with associated intent used to… See the full description on the dataset page: https://huggingface.co/datasets/Wikit/nlu-covid.Long_Covid_word_frequency_TFIDF_22_Feb_AprCOVID-19_qa_pairs
Dataset Card for COVID-19_qa_pairs dataset
Dataset Summary
This datasets includes 604 question-answer pairs related to COVID-19 pandemic machine translated in Greek language.
The data is extracted from the official website of WHO.
Data Fields
question: Query question
document: Answer to the question
Bias, Risks, and Limitations
This dataset is the result of machine translation.
Licensing Information
The dataset is licensed under the… See the full description on the dataset page: https://huggingface.co/datasets/panosgriz/COVID-19_qa_pairs.COVID-19-el-corpus
Dataset Card for
Dataset Summary
This corpus contains Greek-language texts about the COVID-19 pandemic including relevant information, FAQs, etc. The texts were collected from official websites (WHO, ECDC, NPHO, covid19.gov.gr) and articles from the greek Wikipedia. Total number of words: 204,748.
Data Fields
Each instance contains:
content: Plain text
id: Instance ID
title: A document title (only in instances related to Wikipedia articles)
Licensing… See the full description on the dataset page: https://huggingface.co/datasets/panosgriz/COVID-19-el-corpus.CovidRetrieval
🔭 Overview
CMIRB: Chinese Medical Information Retrieval Benchmark
CMIRB is a specialized multi-task dataset designed specifically for medical information retrieval. It consists of data collected from various medical online websites, encompassing 5 tasks and 10 datasets, and has practical application scenarios.
Name
Description
Query #Samples
Doc #Samples
MedExamRetrieval
Medical multi-choice exam
697
27,871
DuBaikeRetrieval
Medical search query from BaiDu… See the full description on the dataset page: https://huggingface.co/datasets/CMIRB/CovidRetrieval.Long_Covid_word_frequency_TFIDF_21_Nov_22_JanLong_Covid_word_frequency_TFIDF_21_Jul_OctLong_Covid_word_frequency_TFIDF_22_May_JulCovidQA
