CoolFace
Datasetpublic

FBK-MT/mosel

Dataset Description, Collection, and Source The MOSEL corpus is a multilingual dataset collection including up to 950K hours of open-source speech recordings covering the 24 official languages of the European Union. We collect data by surveying labeled and unlabeled speech corpora under open-source compliant licenses. In particular, MOSEL includes the automatic transcripts of 441k hours of unlabeled speech from VoxPopuli and LibriLight. The data is transcribed using Whisper… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/mosel.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
93likes2.8kdownloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
FBK-MT/mosel · CoolFace