CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MLCommons /peoples_speech Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.audioautomatic-speech-recognition1M<n<10M285 likes37k downloads2y agoHugging Face02MLCommons /unsupervised_peoples_speech Dataset Card for Unsupervised Peoples Speech Dataset Description Dataset Summary The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers. Point of Contact: MLCommons Datasets Discord Dataset Structure This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.audioautomatic-speech-recognition81 likes26k downloads2y agoHugging Face03parler-tts /mls_eng Dataset Card for English MLS Dataset Summary This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.audioautomatic-speech-recognition10M<n<100M40 likes3.4k downloads2y agoHugging Face04parler-tts /mls_eng_10k Dataset Summary This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.audioautomatic-speech-recognition1M<n<10M31 likes1.1k downloads2y agoHugging Face05mllp /LHCP-ASR LHCP-ASR This dataset is another version of the LHCP-ASR corpus, an English speech dataset for narrow-domain ASR benchmarking in high-energy physics. Unlike the original distribution, which includes video, slides and text data, this version focuses entirely on audio-text pairs DESCRIPTION The speech data are 30 hours of LHCP plenary conference talks (2020, 2022) with manual (human) verbatim transcriptions and 205 hours of LHCP conference talks (2020-2022) with automatic… See the full description on the dataset page: https://huggingface.co/datasets/mllp/LHCP-ASR.audioautomatic-speech-recognition100K<n<1M0 likes773 downloads5mo agoHugging Face06thennal /indic_tts_ml Indic TTS Malayalam Speech Corpus The Malayalam subset of Indic TTS Corpus, taken from this Kaggle database. The corpus contains one male and one female speaker, with a 2:1 ratio of samples due to missing files for the female speaker. The license is given in the repository. audiotext-to-speech1K<n<10K6 likes348 downloads4y agoHugging Face07MLRS /masri_synthetic Dataset Card for masri_synthetic Dataset Summary The MASRI-SYNTHETIC is a corpus made out of synthesized speech in Maltese. The text-to-speech (TTS) system utilized to produce the utterances was developed by the Research & Development Department of Crimsonwing p.l.c. The sentences used to create the corpus were extracted from the MLRS Corpus, which is a corpus of written or transcribed Maltese divided into different genres, including: culture, news, academic, religion… See the full description on the dataset page: https://huggingface.co/datasets/MLRS/masri_synthetic.audioautomatic-speech-recognition10K<n<100K2 likes246 downloads2y agoHugging Face08ntt123 /mls-eng-128kb Dataset Card for English MLS Dataset Summary This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/ntt123/mls-eng-128kb.audioautomatic-speech-recognition1M<n<10M0 likes235 downloads1y agoHugging Face09TTS-AGI /mls-enhanced-dacvae Multilingual LibriSpeech converted to DAC VAE latents Source facebook/multilingual_librispeech Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.audioautomatic-speech-recognition100K<n<1M0 likes218 downloads6mo agoHugging Face10mllp /LHCP-ASR-segments LHCP-ASR Segments This dataset is a segment-level distribution derived from mllp/LHCP-ASR (and the original LHCP-ASR repository), an English speech corpus for narrow-domain ASR benchmarking in high-energy particle physics. Unlike previous versions, this repository provides audio directly at the segment level (<30 seconds each) for evaluation, adds talk-level metadata and cleans up transcription tags. Differences from mllp/LHCP-ASR No subsets: Directly formatted… See the full description on the dataset page: https://huggingface.co/datasets/mllp/LHCP-ASR-segments.audioautomatic-speech-recognition10K<n<100K0 likes121 downloads16d agoHugging Face11Goekdeniz-Guelmez /mlx-omni-lora-stt-tts-demoWill be used in the development of the trainer backend of mlx-omni by Neywa Labs. audioautomatic-speech-recognitionn<1K0 likes112 downloads1mo agoHugging Face12trysem /Shrutilipi-ML Shrutilipi-ML Dataset Overview This dataset is a specific subset of the Shrutilipi ASR (Automatic Speech Recognition) corpus, containing only the Malayalam language data. Shrutilipi is a large-scale multilingual speech dataset for Indian languages, originally curated by AI4Bharat. This repository aims to provide a lightweight, language-specific version for researchers and developers focusing on Malayalam speech technology. Dataset Details Source… See the full description on the dataset page: https://huggingface.co/datasets/trysem/Shrutilipi-ML.audioautomatic-speech-recognition100K<n<1M1 likes80 downloads4mo agoHugging Face13thennal /ulca_ml ULCA ASR Dataset Malayalam Speech Corpus The labelled Malayalam speech subcorpus from the larger ULCA ASR Corpus. The speech is taken from news broadcasts, and is largely composed of short soundbites with some longer outliers. audioautomatic-speech-recognition1K<n<10K1 likes35 downloads4y agoHugging Face14ggfox00000 /stt-mls-test-fr MLS — French test split Split test de Multilingual LibriSpeech (MLS), locale fr, empaqueté en Parquet shardé avec audio FLAC embarqué — prêt pour load_dataset. Usage principal : benchmark ASR français (WER / CER) sur livres audio LibriVox (domaine public). Contenu 2426 utterances Audio : FLAC 16 kHz mono (tel que fourni par MLS upstream) Langue : français (fr) Licence : CC-BY-4.0 (héritée de MLS / OpenSLR 94) Durée totale : 10.07 h Colonnes Colonne Type… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-mls-test-fr.audioautomatic-speech-recognition1K<n<10K0 likes30 downloads5mo agoHugging Face15Bretagne /ml_superb_br Description Partie en breton du jeu de données ML-SUPERB 1.0.En pratique nous nous sommes basés sur espnet/ml_superb_hf.Selon nos estimations, ce jeu de données représente 1 h 10 min et 7s. Citation @misc{shi2025mlsuperbmultilingualspeechuniversal, title={ML-SUPERB: Multilingual Speech Universal PERformance Benchmark}, author={Jiatong Shi and Dan Berrebbi and William Chen and Ho-Lam Chung and En-Pei Hu and Wei Ping Huang and Xuankai Chang and Shang-Wen Li and… See the full description on the dataset page: https://huggingface.co/datasets/Bretagne/ml_superb_br.audioautomatic-speech-recognition1K<n<10K0 likes27 downloads1y agoHugging Face16TTS-AGI /mls-dacvae Multilingual LibriSpeech converted to DAC VAE latents Source facebook/multilingual_librispeech Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-dacvae.audioautomatic-speech-recognition10K<n<100K0 likes27 downloads6mo agoHugging Face17vrclc /festvox-iiith-mlaudioautomatic-speech-recognition1K<n<10K0 likes19 downloads3y agoHugging Face18ggfox00000 /stt-mls-test MLS — French test split Split test de Multilingual LibriSpeech (MLS), locale fr, empaqueté en Parquet shardé avec audio FLAC embarqué — prêt pour load_dataset. Usage principal : benchmark ASR français (WER / CER) sur livres audio LibriVox (domaine public). Contenu 2426 utterances Audio : FLAC 16 kHz mono (tel que fourni par MLS upstream) Langue : français (fr) Licence : CC-BY-4.0 (héritée de MLS / OpenSLR 94) Durée totale : 10.07 h Colonnes Colonne Type… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-mls-test.audioautomatic-speech-recognition1K<n<10K0 likes14 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.