CoolFace
Datasetpublic

mteb/tatoeba-bitext-mining

Tatoeba An MTEB dataset Massive Text Embedding Benchmark 1,000 English-aligned sentence pairs for each language based on the Tatoeba corpus Task category t2t Domains Written Reference https://github.com/facebookresearch/LASER/tree/main/data/tatoeba/v1 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["Tatoeba"]) evaluator = mteb.MTEB(task) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tatoeba-bitext-mining.

sourceHugging Facecc-by-2.0updated 7mo agoView on Hugging Face
9likes1.6kdownloads

mteb/tatoeba-bitext-mining · main · files are served by the source, never re-hosted here