slone/nllb-200-10M-sample
Dataset Card for "nllb-200-10M-sample" This is a sample of nearly 10M sentence pairs from the NLLB-200 mined dataset allenai/nllb, scored with the model facebook/blaser-2.0-qe described in the SeamlessM4T paper. The sample is not random; instead, we just took the top n sentence pairs from each translation direction. The number n was computed with the goal of upsamping the directions that contain underrepresented languages. Nevertheless, the 187 languoids (language and script… See the full description on the dataset page: https://huggingface.co/datasets/slone/nllb-200-10M-sample.
Update README.md
Update README.md
Update README.md
Upload README.md with huggingface_hub
Upload data/train-00004-of-00005-c30bd9963644b892.parquet with huggingface_hub
Upload data/train-00003-of-00005-6d65323f77298b0a.parquet with huggingface_hub
Upload data/train-00002-of-00005-805ecb74951d503f.parquet with huggingface_hub
Upload data/train-00001-of-00005-a6534e9c4eca51c9.parquet with huggingface_hub
Upload data/train-00000-of-00005-cf25f10de200dab1.parquet with huggingface_hub
initial commit
