CoolFace
Datasetpublic

softcatala/catalan-youtube-speech

Catalan YouTube Speech Corpus This dataset contains 231,684 short audio clips of spontaneous Catalan speech, automatically extracted from public YouTube videos. Each clip is paired with two independent machine-generated transcription candidates, along with speaker gender, clip timing, and the source video's reuse license. It was built and published by Softcatalà, the volunteer organization behind free/open-source Catalan-language software. Homepage: https://www.softcatala.org/… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/catalan-youtube-speech.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
3likes147downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
softcatala/catalan-youtube-speech · CoolFace