CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01blanchon /parler-tts_mls_eng_10k_snac_token_old Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.tabularautomatic-speech-recognition100K<n<1M1 likes965 downloads2y agoHugging Face02Gong1212 /emilia-tokengated Emilia EN Pocket Mimi continuous latents This gated repository contains the English Emilia training data representation used by the LatentTTS experiments in this project. Audio was encoded offline with the continuous Gaussian Mimi speech VAE used by Pocket TTS. The files are intended to let an authorized researcher reproduce latent-domain training without encoding the source audio again. Access and licensing This is a derived representation of the official Emilia… See the full description on the dataset page: https://huggingface.co/datasets/Gong1212/emilia-token.audiotext-to-speech0 likes30 downloads18d agoHugging Face03distil-whisper /gigaspeech-l-token-idsGigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised training. Around 40,000 hours of transcribed audio is first collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc. A new forced alignment and segmentation pipeline is proposed to create sentence segments suitable for speech recognition training, and to filter out segments with low-quality transcription. For system training, GigaSpeech provides five subsets of different sizes, 10h, 250h, 1000h, 2500h, and 10000h. For our 10,000-hour XL training subset, we cap the word error rate at 4% during the filtering/validation stage, and for all our other smaller training subsets, we cap it at 0%. The DEV and TEST evaluation sets, on the other hand, are re-processed by professional human transcribers to ensure high transcription quality.automatic-speech-recognition0 likes17 downloads3y agoHugging Face04distil-whisper /librispeech_asr-token-idsLibriSpeech is a corpus of approximately 1000 hours of read English speech with sampling rate of 16 kHz, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.87automatic-speech-recognition0 likes10 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.