softcatala/catalan-youtube-speech
Catalan YouTube Speech Corpus This dataset contains 231,684 short audio clips of spontaneous Catalan speech, automatically extracted from public YouTube videos. Each clip is paired with two independent machine-generated transcription candidates, along with speaker gender, clip timing, and the source video's reuse license. It was built and published by Softcatalà, the volunteer organization behind free/open-source Catalan-language software. Homepage: https://www.softcatala.org/… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/catalan-youtube-speech.
This repository belongs to softcatala on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
