CoolFace
Datasetpublic

amine-khelif/mms_ulab_v2

MMS ulab v2 is a a massively multilingual speech dataset that contains 8900 hours of unlabeled speech across 4023 languages. In total, it contains 189 language families. It can be used for language identification, spoken language modelling, or speech representation learning. MMS ulab v2 is a reproduced and extended version of the MMS ulab dataset originally proposed in Scaling Speech Technology to 1000+ Languages, covering more languages and containing more data. This dataset includes the raw… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/mms_ulab_v2.

sourceHugging Facecc-by-nc-sa-4.0updated 6mo agoView on Hugging Face
0likes235downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

amine-khelif/mms_ulab_v2 · main · files are served by the source, never re-hosted here