datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ClimateFEVER_test_top_250_only_w_correct-v2
ClimateFEVERHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
CLIMATE-FEVER is a dataset adopting the FEVER methodology that consists of 1,535 real-world claims regarding climate-change. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
https://www.sustainablefinance.uzh.ch/en/research/climate-fever.html… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ClimateFEVER_test_top_250_only_w_correct-v2.DBPedia_test_top_250_only_w_correct-v2
DBPediaHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
DBpedia-Entity is a standard test collection for entity search over the DBpedia knowledge base. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Written, Encyclopaedic
Reference
https://github.com/iai-group/DBpedia-Entity/
How to evaluate on this task
You can evaluate… See the full description on the dataset page: https://huggingface.co/datasets/mteb/DBPedia_test_top_250_only_w_correct-v2.HotpotQA_test_top_250_only_w_correct-v2
HotpotQAHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
HotpotQA is a question answering dataset featuring natural, multi-hop questions, with strong supervision for supporting facts to enable more explainable question answering systems. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Web, Written
Reference
https://hotpotqa.github.io/… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HotpotQA_test_top_250_only_w_correct-v2.FEVER_test_top_250_only_w_correct-v2
FEVERHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
FEVER (Fact Extraction and VERification) consists of 185,445 claims generated by altering sentences extracted from Wikipedia and subsequently verified without knowledge of the sentence they were derived from. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Encyclopaedic, Written… See the full description on the dataset page: https://huggingface.co/datasets/mteb/FEVER_test_top_250_only_w_correct-v2.NQ_test_top_250_only_w_correct-v2
NQHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
NFCorpus: A Full-Text Learning to Rank Dataset for Medical Information Retrieval. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
None
Reference
https://ai.google.com/research/NaturalQuestions/
How to evaluate on this task
You can evaluate an embedding model on this… See the full description on the dataset page: https://huggingface.co/datasets/mteb/NQ_test_top_250_only_w_correct-v2.QuoraRetrieval_test_top_250_only_w_correct-v2
QuoraRetrievalHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
QuoraRetrieval is based on questions that are marked as duplicates on the Quora platform. Given a question, find other (duplicate) questions. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
None
Reference
https://quoradata.quora.com/First-Quora-Dataset-Release-Question-Pairs… See the full description on the dataset page: https://huggingface.co/datasets/mteb/QuoraRetrieval_test_top_250_only_w_correct-v2.MUSAN-noise-audio-only
Dataset Card for "MUSAN-noise"
More Information needed
MSMARCO_test_top_250_only_w_correct-v2
MSMARCOHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
MS MARCO is a collection of datasets focused on deep learning in search. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Encyclopaedic, Academic, Blog, News, Medical, Government, Reviews, Non-fiction, Social, Web
Reference
https://microsoft.github.io/msmarco/
How to… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MSMARCO_test_top_250_only_w_correct-v2.RiaNewsRetrieval_test_top_250_only_w_correct-v2
RiaNewsRetrievalHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
News article retrieval by headline. Based on Rossiya Segodnya dataset. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
News, Written
Reference
https://arxiv.org/abs/1901.07786
How to evaluate on this task
You can evaluate an embedding model on this… See the full description on the dataset page: https://huggingface.co/datasets/mteb/RiaNewsRetrieval_test_top_250_only_w_correct-v2.AndroidControlParsedWithImages-20k-TESTONLYmydataset-only-test8B-reason-only.stride-32-test.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
8B-reason-only.stride-32-test.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
Tokenized, tag-wrapped form of JackHsieh/8B-reason-only.stride-32-test.k-8.statml-arxiv.qwen3-ids.
Each thought is wrapped as
<|note|>
This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next.
KEY: <last 8 prefix tokens>
VALUE: <thought>
<|/note|>
and stored both as text (thought_text) and as… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/8B-reason-only.stride-32-test.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.openai-prm800k-phase1_test-solutions-only8B-reason-only.stride-32-test.k-8.statml-arxiv.qwen3-idsopenai-prm800k-phase2_test-solutions-onlyMIMIC-Impression-Only-Data-Train-TestMSMARCO_PL_test_top_250_only_w_correct-v2
MSMARCO-PLHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
MS MARCO is a collection of datasets focused on deep learning in search. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Web, Written
Reference
https://microsoft.github.io/msmarco/
How to evaluate on this task
You can evaluate an embedding model on this dataset… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MSMARCO_PL_test_top_250_only_w_correct-v2.NQ_PL_test_top_250_only_w_correct-v2
NQ-PLHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
Natural Questions: A Benchmark for Question Answering Research. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task categoryt2t
Domains
None
Reference
https://ai.google.com/research/NaturalQuestions/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/mteb/NQ_PL_test_top_250_only_w_correct-v2.test_math_shepherd_only_solutions
Dataset Card for test_math_shepherd_only_solutions
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/plaguss/test_math_shepherd_only_solutions/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/plaguss/test_math_shepherd_only_solutions.Thai-Voice-Test-Speaker-Only
Thanarit/Thai-Voice
Combined Thai audio dataset from multiple sources
Dataset Details
Total samples: 20
Total duration: 0.02 hours
Language: Thai (th)
Audio format: 16kHz mono WAV
Volume normalization: -20dB
Sources
Processed 1 datasets in streaming mode
Source Datasets
GigaSpeech2: Large-scale multilingual speech corpus
Usage
from datasets import load_dataset
# Load with streaming to avoid downloading everything
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Thanarit/Thai-Voice-Test-Speaker-Only.Quora_PL_test_top_250_only_w_correct-v2
Quora-PLHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
QuoraRetrieval is based on questions that are marked as duplicates on the Quora platform. Given a question, find other (duplicate) questions. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
None
Reference
https://quoradata.quora.com/First-Quora-Dataset-Release-Question-Pairs… See the full description on the dataset page: https://huggingface.co/datasets/mteb/Quora_PL_test_top_250_only_w_correct-v2.NQ_PL_test_top_250_only_w_correctDBPedia_PL_test_top_250_only_w_correct-v2
DBPedia-PLHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
DBpedia-Entity is a standard test collection for entity search over the DBpedia knowledge base. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Written, Encyclopaedic
Reference
https://github.com/iai-group/DBpedia-Entity/
How to evaluate on this task
You can… See the full description on the dataset page: https://huggingface.co/datasets/mteb/DBPedia_PL_test_top_250_only_w_correct-v2.HotpotQA_PL_test_top_250_only_w_correct-v2
HotpotQA-PLHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
HotpotQA is a question answering dataset featuring natural, multi-hop questions, with strong supervision for supporting facts to enable more explainable question answering systems. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Web, Written
Reference… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HotpotQA_PL_test_top_250_only_w_correct-v2.DBPedia_PL_test_top_250_only_w_correctNQ_test_top_250_only_w_correctonly-test-magento2MSMARCO_test_top_250_only_w_correcttestonly
Dataset Card for "testonly"
More Information needed
RiaNewsRetrieval_test_top_250_only_w_correct
