test-only
ClimateFEVER_test_top_250_only_w_correct-v2
ClimateFEVERHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
CLIMATE-FEVER is a dataset adopting the FEVER methodology that consists of 1,535 real-world claims regarding climate-change. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
https://www.sustainablefinance.uzh.ch/en/research/climate-fever.html… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ClimateFEVER_test_top_250_only_w_correct-v2.DBPedia_test_top_250_only_w_correct-v2
DBPediaHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
DBpedia-Entity is a standard test collection for entity search over the DBpedia knowledge base. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Written, Encyclopaedic
Reference
https://github.com/iai-group/DBpedia-Entity/
How to evaluate on this task
You can evaluate… See the full description on the dataset page: https://huggingface.co/datasets/mteb/DBPedia_test_top_250_only_w_correct-v2.HotpotQA_test_top_250_only_w_correct-v2
HotpotQAHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
HotpotQA is a question answering dataset featuring natural, multi-hop questions, with strong supervision for supporting facts to enable more explainable question answering systems. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Web, Written
Reference
https://hotpotqa.github.io/… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HotpotQA_test_top_250_only_w_correct-v2.FEVER_test_top_250_only_w_correct-v2
FEVERHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
FEVER (Fact Extraction and VERification) consists of 185,445 claims generated by altering sentences extracted from Wikipedia and subsequently verified without knowledge of the sentence they were derived from. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Encyclopaedic, Written… See the full description on the dataset page: https://huggingface.co/datasets/mteb/FEVER_test_top_250_only_w_correct-v2.NQ_test_top_250_only_w_correct-v2
NQHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
NFCorpus: A Full-Text Learning to Rank Dataset for Medical Information Retrieval. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
None
Reference
https://ai.google.com/research/NaturalQuestions/
How to evaluate on this task
You can evaluate an embedding model on this… See the full description on the dataset page: https://huggingface.co/datasets/mteb/NQ_test_top_250_only_w_correct-v2.QuoraRetrieval_test_top_250_only_w_correct-v2
QuoraRetrievalHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
QuoraRetrieval is based on questions that are marked as duplicates on the Quora platform. Given a question, find other (duplicate) questions. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
None
Reference
https://quoradata.quora.com/First-Quora-Dataset-Release-Question-Pairs… See the full description on the dataset page: https://huggingface.co/datasets/mteb/QuoraRetrieval_test_top_250_only_w_correct-v2.
