styletts2
multilingual-pl-bertAttribution: Wikipedia.org
multilingual-phonemes-10k-alpha
Multilingual Phonemes 10K Alpha
This dataset contains approximately 10,000 pairs of text and phonemes from each supported language. We support 15 languages in this dataset, so we have a total of ~150K pairs. This does not include the English-XL dataset, which includes another 100K unique rows.
Languages
We support 15 languages, which means we have around 150,000 pairs of text and phonemes in multiple languages. This excludes the English-XL dataset, which has 100K unique… See the full description on the dataset page: https://huggingface.co/datasets/styletts2-community/multilingual-phonemes-10k-alpha.Common_voice_French_StyleTTS2_63_speakerscommon-voice-filtered
Common Voice Filtered
A filtered subset of the Common Voice dataset. Currently, this dataset only includes a small subset of English speech.
We only include speech ranked above 3.75 (75%) on the MOS metric, as calculated by the UTMOS system. Approximately 7% of audio qualified for inclusion in this filtered dataset.
This data is not final. Processing the whole Common Voice dataset would require a significant amount of compute, this is just a small sample/MVP of the project.
The code… See the full description on the dataset page: https://huggingface.co/datasets/styletts2-community/common-voice-filtered.styletts2_datasetStyleTTS2MLP
