CoolFace
Datasetpublic

SaylorTwift/mteb-bitext-mining-aggregated

MTEB BitextMining Aggregated Dataset (Full) This dataset aggregates ALL configs from 10 BitextMining datasets in the MTEB (Massive Text Embedding Benchmark) Multilingual v2 benchmark into a single, unified dataset for comprehensive bitext mining evaluation. Dataset Summary Total Examples: 448,229 sentence pairs Source Datasets (Configs): 10 MTEB BitextMining tasks Total Splits: 332 language pairs/configurations Languages: 300+ unique language codes across all… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/mteb-bitext-mining-aggregated.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes896downloads
settings

This repository belongs to SaylorTwift on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namemteb-bitext-mining-aggregated
visibilitypublic
licenceapache-2.0
gatedno
ownerSaylorTwift
Account settings
SaylorTwift/mteb-bitext-mining-aggregated · CoolFace