datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
libritts_r_filtered
Dataset Card for Filtered LibriTTS-R
This is a filtered version of LibriTTS-R. It has been filtered based on two sources:
LibriTTS-R paper [1], which lists samples for which speech restoration have failed
LibriTTS-P [2] list of excluded speakers for which multiple speakers have been detected.
LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately
585 hours of read English speech at 24kHz sampling rate… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts_r_filtered.mls_eng
Dataset Card for English MLS
Dataset Summary
This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.mls_eng_10k
Dataset Summary
This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.parlertts-pony-speech-audiohassaniya-parler-ttsparlertts-pony-speech-audioparler-tts-dataset-balancedparler-tts-emotion-datasetParler-TTS-Datadriven-100h-44.1kHz_stage1parler-tts-mini-v1-a_speaker_similarityparler-tts-large_nve_samplesparler-tts-mini-v1_speaker_similarityparler-tts-large-v1-wsd_speaker_similaritytest_parler_tts_train_dataIndic_parler_tts_testing_datasetparler-tts-mini-v1-fast_speaker_similarityparler-tts-dataset-mergednepali-tts-amitpant7-parlerparler_tts_mini_V01_TestVoice_ItalianIndic_parler_tts_audio_datasetparler-tts-large-v1_speaker_similarityparler-tts-large-v1-a_speaker_similarityparler_tts_mini_V01_TestVoice_Italian_V1haruhi-parler-tts-v1custom_parlertts_datasetharuhi-parler-tts-tagged-v1
