datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trec-covid
TRECCOVID
An MTEB dataset
Massive Text Embedding Benchmark
TRECCOVID is an ad-hoc search challenge based on the COVID-19 dataset containing scientific articles related to the COVID-19 pandemic.
Task category
t2t
Domains
Medical, Academic, Written
Reference
https://ir.nist.gov/covidSubmit/index.html
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TRECCOVID"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/trec-covid.TREC-QC
TREC Question Classification
Question classification in coarse and fine-grained categories.
Source:
Experimental Data for Question Classification
Xin Li, Dan Roth, Learning Question Classifiers. COLING'02, Aug., 2002.
trec-covid
Dataset Card for BEIR Benchmark
trec-covid is one of the datasets from the Bio-Medical Retrieval task within BEIR, measuring scientific article retrieval for a given query on COVID-19.
Dataset Summary
BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks.
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/trec-covid.trec-covid-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/trec-covid-qrels.AToMiC-Images-v0.2
Dataset Card for "AToMiC-All-Images_wi-pixels"
Languages
The dataset contains 108 languages in Wikipedia.
Data Instances
Each instance is an image, its representation in bytes, and its associated captions.
Intended Usage
Image collection for Text-to-Image retrieval
Image--Caption Retrieval/Generation/Translation
Licensing Information
CC BY-SA 4.0 international license
Citation Information
TBA
Acknowledgement
Thanks… See the full description on the dataset page: https://huggingface.co/datasets/TREC-AToMiC/AToMiC-Images-v0.2.AToMiC-Baselines
AToMiC Prebuilt Indexes
Example Usage:
Reproduction
Toolkits:
https://github.com/TREC-AToMiC/AToMiC/tree/main/examples/dense_retriever_baselines
# Skip the encode and index steps, search with the prebuilt indexes and topics directly
python search.py \
--topics topics/openai.clip-vit-base-patch32.text.validation \
--index indexes/openai.clip-vit-base-patch32.image.faiss.flat \
--hits 1000 \
--output… See the full description on the dataset page: https://huggingface.co/datasets/TREC-AToMiC/AToMiC-Baselines.trec-covid-decontaminated
trec-covid (Decontaminated)
A decontaminated version of the trec-covid dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed.
Decontamination methodology
Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files):
Pass 1: Exact hash matching
All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed with… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/trec-covid-decontaminated.trec-ragtime-2026
TREC RAGTIME 2026 — sentence and passage renderings
A sentence-level view of the TREC RAGTIME 2026 news collection,
with two English machine translations of every non-English sentence and the passage boundaries used
for retrieval. Derived from trec-ragtime/ragtime2.
Pipeline, experiment design, run configurations and reproduction steps:
github.com/jknafou/trec-ragtime-2026
What is in here
Config
Splits
Rows
Contents
sentences
eng, spa, rus, zho
88,719… See the full description on the dataset page: https://huggingface.co/datasets/jknafou/trec-ragtime-2026.ragtime1
RAGTIME1 Collection
This dataset contains the documents for TREC RAGTIME Track.
Please refer to the website for the details of the task.
RAGTIME is a multilingual RAG task, which expects the participating system to retrieve relevant documents from all four languages and synthesize a response with citation to the report request.
For convenience, we separate the documents by their languages into four .jsonl files. However, they are intended to be used as a whole set.
The documents… See the full description on the dataset page: https://huggingface.co/datasets/trec-ragtime/ragtime1.trec6trec_ragliveqa_medical_trec2017
Dataset Card for LiveQA Medical from TREC 2017
The LiveQA'17 medical task focuses on consumer health question answering. Consumer health questions were received by the U.S. National Library of Medicine (NLM).
The dataset consists of constructed medical question-answer pairs for training and testing, with additional annotations that can be used to develop question analysis and question answering systems.
Please refer to our overview paper for more information about the constructed… See the full description on the dataset page: https://huggingface.co/datasets/hyesunyun/liveqa_medical_trec2017.trec
TREC
A classic benchmark dataset for question classification with both coarse and fine-grained labels.
Size: small, clean, ready to use
Source: original release
Format: stored in Parquet
Compatibility: 🧩 works with datasets >= 4.0 (script loaders deprecated)
Reference
Li, X., & Roth, D. (2002).Learning Question Classifiers.ACL Anthology
trecqa
Dataset Card for "trecqa"
TREC-QA dataset for Answer Sentence Selection. The dataset contains 2 additional splits which are clean versions of the original development and test sets. clean versions contain only questions which have at least a positive and a negative answer candidate.
merged-trecdltrec-dl-2019
TRECDL2019
An MTEB dataset
Massive Text Embedding Benchmark
TREC Deep Learning Track 2019 passage ranking task. The task involves retrieving relevant passages from the MS MARCO collection given web search queries. Queries have multi-graded relevance judgments.
Task categoryt2t
Domains
Encyclopaedic, Academic, Blog, News, Medical, Government, Reviews, Non-fiction, Social, Web
Reference
https://microsoft.github.io/msmarco/TREC-Deep-Learning-2019
How to… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/trec-dl-2019.TREC_Clinical-Trials
TREC Clinical Trials (2021, 2022 and 2023)
Dataset Description
Links
Homepage:
TREC
Paper:
2021 / 2022 / 2023
Contact (Original Authors):
-
Contact (Curator):
Artur Guimarães (artur.guimas@gmail.com)
Dataset Summary
This track focused on matching patients to relevant clinical trials.
Data Instances
Source Format
{
"query_id":"1",
"disease":"glaucoma",
"text":"definitive diagnosis: primary open angle… See the full description on the dataset page: https://huggingface.co/datasets/araag2/TREC_Clinical-Trials.trec-newsclinical-trials-trec-qrelsclinical-trials-trec-parsedatomic2024trec6clinical-trials-trec-topicsproduct-recommendation-2025
TREC 2025 Product Recommendation Data
This is the data for the recommendation task for the TREC 2025 Product Search and Recommendation Task.
The initial directory contains the initial corpus and training data release.
This may be updated as we get further along in the timeline.
[!NOTE]
This data is derived from the Amazon ESCI and M2 data sets, each under the
Apache license (version 2.0).
Product-Search-Images-v0.1TRECCOVID-PL
TRECCOVID-PL
An MTEB dataset
Massive Text Embedding Benchmark
TRECCOVID is an ad-hoc search challenge based on the COVID-19 dataset containing scientific articles related to the COVID-19 pandemic.
Task category
t2t
Domains
Academic, Medical, Non-fiction, Written
Reference
https://ir.nist.gov/covidSubmit/index.html
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/TRECCOVID-PL.benchmark-trec-covidtrec-covid-vn
TRECCOVID-VN
An MTEB dataset
Massive Text Embedding Benchmark
A translated dataset from TRECCOVID is an ad-hoc search challenge based on the COVID-19 dataset containing scientific articles related to the COVID-19 pandemic. The process of creating the VN-MTEB (Vietnamese Massive Text Embedding Benchmark) from English samples involves a new automated system: - The system uses large language models (LLMs), specifically Coherence's Aya model, for translation. - Applies advanced embedding… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/trec-covid-vn.trecbeir-trec-covid
TRECCOVID — BEIR, unified schema
A normalised copy of the dataset behind the mteb task TRECCOVID, one of the tasks of the BEIR benchmark as mteb defines it. Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/trec-covid @ bb9466bac815 (the revision pinned in mteb)
Domain · languages
biomedical · eng
Queries / documents / qrels
50 / 171,332 / 66,336… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-trec-covid.
