CoolFace
Datasetpublic

slone/nllb-200-10M-sample

Dataset Card for "nllb-200-10M-sample" This is a sample of nearly 10M sentence pairs from the NLLB-200 mined dataset allenai/nllb, scored with the model facebook/blaser-2.0-qe described in the SeamlessM4T paper. The sample is not random; instead, we just took the top n sentence pairs from each translation direction. The number n was computed with the goal of upsamping the directions that contain underrepresented languages. Nevertheless, the 187 languoids (language and script… See the full description on the dataset page: https://huggingface.co/datasets/slone/nllb-200-10M-sample.

sourceHugging Faceodc-byupdated 3y agoView on Hugging Face
14likes334downloads
settings

This repository belongs to slone on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namenllb-200-10M-sample
visibilitypublic
licenceodc-by
gatedno
ownerslone
Account settings
slone/nllb-200-10M-sample · CoolFace