CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MLCommons /peoples_speech Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.audioautomatic-speech-recognition1M<n<10M285 likes39k downloads2y agoHugging Face02MLCommons /unsupervised_peoples_speech Dataset Card for Unsupervised Peoples Speech Dataset Description Dataset Summary The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers. Point of Contact: MLCommons Datasets Discord Dataset Structure This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.audioautomatic-speech-recognition81 likes26k downloads2y agoHugging Face03MLCommons /peoples_speech_v1.0 Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech_v1.0.automatic-speech-recognition8 likes6.9k downloads2y agoHugging Face04parler-tts /mls_eng Dataset Card for English MLS Dataset Summary This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.audioautomatic-speech-recognition10M<n<100M40 likes3.5k downloads2y agoHugging Face05parler-tts /mls_eng_10k Dataset Summary This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.audioautomatic-speech-recognition1M<n<10M31 likes1.1k downloads2y agoHugging Face06blanchon /parler-tts_mls_eng_10k_snac_token_old Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.tabularautomatic-speech-recognition100K<n<1M1 likes991 downloads2y agoHugging Face07mllp /LHCP-ASR LHCP-ASR This dataset is another version of the LHCP-ASR corpus, an English speech dataset for narrow-domain ASR benchmarking in high-energy physics. Unlike the original distribution, which includes video, slides and text data, this version focuses entirely on audio-text pairs DESCRIPTION The speech data are 30 hours of LHCP plenary conference talks (2020, 2022) with manual (human) verbatim transcriptions and 205 hours of LHCP conference talks (2020-2022) with automatic… See the full description on the dataset page: https://huggingface.co/datasets/mllp/LHCP-ASR.audioautomatic-speech-recognition100K<n<1M0 likes772 downloads5mo agoHugging Face08parler-tts /mls-eng-speaker-descriptions Dataset Card for Annotations of English MLS This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-speaker-descriptions.tabularautomatic-speech-recognition10M<n<100M13 likes400 downloads2y agoHugging Face09thennal /indic_tts_ml Indic TTS Malayalam Speech Corpus The Malayalam subset of Indic TTS Corpus, taken from this Kaggle database. The corpus contains one male and one female speaker, with a 2:1 ratio of samples due to missing files for the female speaker. The license is given in the repository. audiotext-to-speech1K<n<10K6 likes343 downloads4y agoHugging Face10MLRS /masri_synthetic Dataset Card for masri_synthetic Dataset Summary The MASRI-SYNTHETIC is a corpus made out of synthesized speech in Maltese. The text-to-speech (TTS) system utilized to produce the utterances was developed by the Research & Development Department of Crimsonwing p.l.c. The sentences used to create the corpus were extracted from the MLRS Corpus, which is a corpus of written or transcribed Maltese divided into different genres, including: culture, news, academic, religion… See the full description on the dataset page: https://huggingface.co/datasets/MLRS/masri_synthetic.audioautomatic-speech-recognition10K<n<100K2 likes260 downloads2y agoHugging Face11bsmu /MLC-SLM-Eval Interspeech2025 Multilingual Conversational Speech Language Model (MLC-SLM) Eval Groundtruth 🖥️ Overview In the MLC-SLM challenge, we only provided the participants with the audio files of the Eval sets. Now, we release the oracle segmentation, speaker labels, and transcriptions of the Eval sets to facilitate further research by all participants on the MLC-SLM dataset! In addition, the MLC-SLM challenge summary paper "Summary on The Multilingual Conversational Speech… See the full description on the dataset page: https://huggingface.co/datasets/bsmu/MLC-SLM-Eval.automatic-speech-recognition10K<n<100K3 likes248 downloads1y agoHugging Face12ntt123 /mls-eng-128kb Dataset Card for English MLS Dataset Summary This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/ntt123/mls-eng-128kb.audioautomatic-speech-recognition1M<n<10M0 likes234 downloads1y agoHugging Face13TTS-AGI /mls-enhanced-dacvae Multilingual LibriSpeech converted to DAC VAE latents Source facebook/multilingual_librispeech Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.audioautomatic-speech-recognition100K<n<1M0 likes214 downloads6mo agoHugging Face14parler-tts /mls-eng-10k-tags_tagged_10k_generated Dataset Card for Annotations of 10K hours of English MLS This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-10k-tags_tagged_10k_generated.tabularautomatic-speech-recognition1M<n<10M17 likes138 downloads2y agoHugging Face15mllp /LHCP-ASR-segments LHCP-ASR Segments This dataset is a segment-level distribution derived from mllp/LHCP-ASR (and the original LHCP-ASR repository), an English speech corpus for narrow-domain ASR benchmarking in high-energy particle physics. Unlike previous versions, this repository provides audio directly at the segment level (<30 seconds each) for evaluation, adds talk-level metadata and cleans up transcription tags. Differences from mllp/LHCP-ASR No subsets: Directly formatted… See the full description on the dataset page: https://huggingface.co/datasets/mllp/LHCP-ASR-segments.audioautomatic-speech-recognition10K<n<100K0 likes120 downloads15d agoHugging Face16Goekdeniz-Guelmez /mlx-omni-lora-stt-tts-demoWill be used in the development of the trainer backend of mlx-omni by Neywa Labs. audioautomatic-speech-recognitionn<1K0 likes115 downloads1mo agoHugging Face17nccm2p2 /MLD-VC 🎥 MLD-VC: Multimodal Dataset for Video Conferencing When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse (CVPR 2026) 📄 [Paper] | 🤗 [Hugging Face Dataset] 📌 Overview MLD-VC is the first multimodal dataset specifically designed for Audio-Visual Speech Recognition (AVSR) in real-world video conferencing (VC) scenarios. Unlike traditional AVSR datasets collected in controlled offline environments, MLD-VC… See the full description on the dataset page: https://huggingface.co/datasets/nccm2p2/MLD-VC.automatic-speech-recognition0 likes83 downloads6mo agoHugging Face18trysem /Shrutilipi-ML Shrutilipi-ML Dataset Overview This dataset is a specific subset of the Shrutilipi ASR (Automatic Speech Recognition) corpus, containing only the Malayalam language data. Shrutilipi is a large-scale multilingual speech dataset for Indian languages, originally curated by AI4Bharat. This repository aims to provide a lightweight, language-specific version for researchers and developers focusing on Malayalam speech technology. Dataset Details Source… See the full description on the dataset page: https://huggingface.co/datasets/trysem/Shrutilipi-ML.audioautomatic-speech-recognition100K<n<1M1 likes79 downloads4mo agoHugging Face19pharaouk /mls-eng-10k-tags_tagged_10k_generated Dataset Card for Annotations of 10K hours of English MLS This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/pharaouk/mls-eng-10k-tags_tagged_10k_generated.tabularautomatic-speech-recognition1M<n<10M0 likes61 downloads2y agoHugging Face20shangeth /mls-mimi-codes Multilingual LibriSpeech (MLS) — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for Multilingual LibriSpeech — LibriVox audiobooks in 7 non-English languages. English is intentionally excluded. For English Mimi codes, use: shangeth/librispeech-mimi-codes — LibriSpeech (~280k rows, 7 splits) shangeth/libritts-r-mimi-codes — LibriTTS-R (~360k rows, 7 splits, 24 kHz native) shangeth/vctk-mimi-codes — VCTK (~44k rows, 110 speakers w/ accents) shangeth/jenny-mimi-codes — Jenny… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/mls-mimi-codes.tabulartext-to-speech1M<n<10M0 likes48 downloads5mo agoHugging Face21Whispering-GPT /whisper-transcripts-ml-street-talk Dataset Card for "whisper-transcripts-mlst" More Information needed textautomatic-speech-recognitionn<1K1 likes41 downloads4y agoHugging Face22thennal /ulca_ml ULCA ASR Dataset Malayalam Speech Corpus The labelled Malayalam speech subcorpus from the larger ULCA ASR Corpus. The speech is taken from news broadcasts, and is largely composed of short soundbites with some longer outliers. audioautomatic-speech-recognition1K<n<10K1 likes39 downloads4y agoHugging Face23Bretagne /ml_superb_br Description Partie en breton du jeu de données ML-SUPERB 1.0.En pratique nous nous sommes basés sur espnet/ml_superb_hf.Selon nos estimations, ce jeu de données représente 1 h 10 min et 7s. Citation @misc{shi2025mlsuperbmultilingualspeechuniversal, title={ML-SUPERB: Multilingual Speech Universal PERformance Benchmark}, author={Jiatong Shi and Dan Berrebbi and William Chen and Ho-Lam Chung and En-Pei Hu and Wei Ping Huang and Xuankai Chang and Shang-Wen Li and… See the full description on the dataset page: https://huggingface.co/datasets/Bretagne/ml_superb_br.audioautomatic-speech-recognition1K<n<10K0 likes29 downloads1y agoHugging Face24ggfox00000 /stt-mls-test-fr MLS — French test split Split test de Multilingual LibriSpeech (MLS), locale fr, empaqueté en Parquet shardé avec audio FLAC embarqué — prêt pour load_dataset. Usage principal : benchmark ASR français (WER / CER) sur livres audio LibriVox (domaine public). Contenu 2426 utterances Audio : FLAC 16 kHz mono (tel que fourni par MLS upstream) Langue : français (fr) Licence : CC-BY-4.0 (héritée de MLS / OpenSLR 94) Durée totale : 10.07 h Colonnes Colonne Type… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-mls-test-fr.audioautomatic-speech-recognition1K<n<10K0 likes29 downloads5mo agoHugging Face25TTS-AGI /mls-dacvae Multilingual LibriSpeech converted to DAC VAE latents Source facebook/multilingual_librispeech Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-dacvae.audioautomatic-speech-recognition10K<n<100K0 likes27 downloads6mo agoHugging Face26vrclc /festvox-iiith-mlaudioautomatic-speech-recognition1K<n<10K0 likes19 downloads3y agoHugging Face27sandylolpotty /MLDSUM_NEWtexttext-classificationn<1K0 likes15 downloads1y agoHugging Face28juancopi81 /mls Dataset Card for "mls" More Information needed textautomatic-speech-recognitionn<1K0 likes14 downloads4y agoHugging Face29french-datasets /espnet_ml_superb_hfCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données espnet/ml_superb_hf. automatic-speech-recognition0 likes13 downloads1y agoHugging Face30ggfox00000 /stt-mls-test MLS — French test split Split test de Multilingual LibriSpeech (MLS), locale fr, empaqueté en Parquet shardé avec audio FLAC embarqué — prêt pour load_dataset. Usage principal : benchmark ASR français (WER / CER) sur livres audio LibriVox (domaine public). Contenu 2426 utterances Audio : FLAC 16 kHz mono (tel que fourni par MLS upstream) Langue : français (fr) Licence : CC-BY-4.0 (héritée de MLS / OpenSLR 94) Durée totale : 10.07 h Colonnes Colonne Type… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-mls-test.audioautomatic-speech-recognition1K<n<10K0 likes13 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.