datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-pl-bertAttribution: Wikipedia.org
multilingual-phonemes-10k-alpha
Multilingual Phonemes 10K Alpha
This dataset contains approximately 10,000 pairs of text and phonemes from each supported language. We support 15 languages in this dataset, so we have a total of ~150K pairs. This does not include the English-XL dataset, which includes another 100K unique rows.
Languages
We support 15 languages, which means we have around 150,000 pairs of text and phonemes in multiple languages. This excludes the English-XL dataset, which has 100K unique… See the full description on the dataset page: https://huggingface.co/datasets/styletts2-community/multilingual-phonemes-10k-alpha.Common_voice_French_StyleTTS2_63_speakersstyletts2_datasetstyletts2-plt-corpus
Mimba StyleTTS2 PLT Corpus — Plateau Malagasy TTS Training Corpus
A ready-to-train corpus for StyleTTS2 on Plateau Malagasy (PLT), pairing
clean audio, original text, and IPA-phonemized text for one or more reference
speakers. Built as a unified, self-contained HuggingFace dataset so that training
notebooks can load a single source of truth — audio, transcription and speaker
metadata in one place — without juggling multiple files or repositories.
⚠️ Derived from synthetic… See the full description on the dataset page: https://huggingface.co/datasets/mimba/styletts2-plt-corpus.styletts2-plt-filelists
