CoolFace
Datasetpublic

slone/nllb-200-10M-sample

Dataset Card for "nllb-200-10M-sample" This is a sample of nearly 10M sentence pairs from the NLLB-200 mined dataset allenai/nllb, scored with the model facebook/blaser-2.0-qe described in the SeamlessM4T paper. The sample is not random; instead, we just took the top n sentence pairs from each translation direction. The number n was computed with the goal of upsamping the directions that contain underrepresented languages. Nevertheless, the 187 languoids (language and script… See the full description on the dataset page: https://huggingface.co/datasets/slone/nllb-200-10M-sample.

sourceHugging Faceodc-byupdated 3y agoView on Hugging Face
14likes335downloads
10 commits on main
112cd773y ago

Update README.md

cointegrated
5e91e6c3y ago

Update README.md

cointegrated
fe7b8143y ago

Update README.md

cointegrated
fb9c5873y ago

Upload README.md with huggingface_hub

cointegrated
99b79773y ago

Upload data/train-00004-of-00005-c30bd9963644b892.parquet with huggingface_hub

cointegrated
a7c77af3y ago

Upload data/train-00003-of-00005-6d65323f77298b0a.parquet with huggingface_hub

cointegrated
086619f3y ago

Upload data/train-00002-of-00005-805ecb74951d503f.parquet with huggingface_hub

cointegrated
86aa52f3y ago

Upload data/train-00001-of-00005-a6534e9c4eca51c9.parquet with huggingface_hub

cointegrated
679e2b53y ago

Upload data/train-00000-of-00005-cf25f10de200dab1.parquet with huggingface_hub

cointegrated
e5066b53y ago

initial commit

cointegrated