CoolFace
Datasetpublic

amine-khelif/mms_ulab_v2

MMS ulab v2 is a a massively multilingual speech dataset that contains 8900 hours of unlabeled speech across 4023 languages. In total, it contains 189 language families. It can be used for language identification, spoken language modelling, or speech representation learning. MMS ulab v2 is a reproduced and extended version of the MMS ulab dataset originally proposed in Scaling Speech Technology to 1000+ Languages, covering more languages and containing more data. This dataset includes the raw… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/mms_ulab_v2.

sourceHugging Facecc-by-nc-sa-4.0updated 6mo agoView on Hugging Face
0likes235downloads
settings

This repository belongs to amine-khelif on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namemms_ulab_v2
visibilitypublic
licencecc-by-nc-sa-4.0
gatedno
owneramine-khelif
Account settings
amine-khelif/mms_ulab_v2 · CoolFace