datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.bangla-10k
Bangla-10K: A Challenging, Metadata-Rich Corpus of Read and Conversational Bengali Speech from India and Bangladesh
Bangla-10K is a 10,816-hour Bengali speech corpus with
624,951 recordings from India and Bangladesh: a 10,070.8-hour core corpus
(567,323 recordings) and a separately collected 745.1-hour evaluation set
(57,628 recordings). It combines scripted single-speaker read speech with
natural multi-speaker conversations for Bengali automatic speech recognition
(ASR).
The… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bangla-10k.mls-eng-10k-tags_tagged_10k_generated
Dataset Card for Annotations of 10K hours of English MLS
This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-10k-tags_tagged_10k_generated.mls-eng-10k-tags_tagged_10k_generated
Dataset Card for Annotations of 10K hours of English MLS
This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/pharaouk/mls-eng-10k-tags_tagged_10k_generated.
