datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
talromur-rosa
Talrómur Dataset (ROSA Speaker Subset)
This Huggingface dataset card describes a subset of the Talrómur speech corpus, focusing exclusively on the "ROSA" speaker data. This is not my original dataset, but rather a subset of the publicly available Talrómur corpus.
Dataset Description
Talrómur is a public domain speech corpus designed for text-to-speech research and development. The complete corpus consists of 122,417 short audio clips from eight different speakers reading… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/talromur-rosa.catalan-datasetrosa_braw_starssamuel_rosa
