CoolFace
Datasetpublic

mteb/TopiOCQA_validation_top_250_only_w_correct-v2

TopiOCQAHardNegatives An MTEB dataset Massive Text Embedding Benchmark TopiOCQA (Human-in-the-loop Attributable Generative Retrieval for Information-seeking Dataset) is information-seeking conversational dataset with challenging topic switching phenomena. It consists of conversation histories along with manually labelled relevant/gold passage. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.… See the full description on the dataset page: https://huggingface.co/datasets/mteb/TopiOCQA_validation_top_250_only_w_correct-v2.

sourceHugging Facecc-by-nc-sa-4.0updated 1y agoView on Hugging Face
0likes29downloads
6 commits on main
2fc419f1y ago

Add dataset card

Samoed
a8339c41y ago

Add dataset card

Samoed
b4cc09f2y ago

Upload dataset

orionweller
9a86ef52y ago

Upload dataset

orionweller
72a0cc62y ago

Upload dataset

orionweller
516b7332y ago

initial commit

orionweller