CoolFace
Datasetpublic

mteb/MIRACLRetrieval

MIRACLRetrieval An MTEB dataset Massive Text Embedding Benchmark MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages. Task category t2t Domains Encyclopaedic, Written Reference http://miracl.ai/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MIRACLRetrieval.

sourceHugging Facecc-by-sa-4.0updated 1y agoView on Hugging Face
9likes2.9kdownloads
../
filedev-00000-of-00012.parquet236.1 MBdownload
filedev-00001-of-00012.parquet265.4 MBdownload
filedev-00002-of-00012.parquet251.1 MBdownload
filedev-00003-of-00012.parquet248.4 MBdownload
filedev-00004-of-00012.parquet242.7 MBdownload
filedev-00005-of-00012.parquet270.0 MBdownload
filedev-00006-of-00012.parquet243.0 MBdownload
filedev-00007-of-00012.parquet242.2 MBdownload
filedev-00008-of-00012.parquet242.2 MBdownload
filedev-00009-of-00012.parquet238.6 MBdownload
filedev-00010-of-00012.parquet241.9 MBdownload
filedev-00011-of-00012.parquet226.0 MBdownload

mteb/MIRACLRetrieval · main · files are served by the source, never re-hosted here