CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.audioautomatic-speech-recognition100K<n<1M468 likes101k downloads4mo agoHugging Face02espnet /yodas-granary Dataset Card for YODAS-Granary Repository: NeMo-speech-data-processor: Granary Paper: Granary: Speech Recognition and Translation Dataset in 25 European Languages Shared by: ESPnet Dataset Description YODAS-Granary is a curated subset of the larger nvidia/Granary dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages. Overview… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas-granary.audioautomatic-speech-recognition10M<n<100M33 likes79k downloads1y agoHugging Face03facebook /voxpopuli Dataset Card for Voxpopuli Dataset Summary VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials. This implementation contains transcribed speech data for 18 languages. It also contains 29 hours of transcribed speech data of non-native… See the full description on the dataset page: https://huggingface.co/datasets/facebook/voxpopuli.audioautomatic-speech-recognition1M<n<10M164 likes77k downloads8mo agoHugging Face04disco-eth /WorldSpeech WorldSpeech A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language-region variants, drawn from national parliaments, public broadcasters, public-domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human-provided transcript, an aligned ASR transcript, character error rate (CER) between the two, a WADA-SNR estimate, and four DNSMOS-P.835 quality scores. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/WorldSpeech.audioautomatic-speech-recognition10M<n<100M49 likes55k downloads4mo agoHugging Face05openslr /librispeech_asr Dataset Card for librispeech_asr Dataset Summary LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned. Supported Tasks and Leaderboards automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/openslr/librispeech_asr.audioautomatic-speech-recognition100K<n<1M245 likes54k downloads1y agoHugging Face06amphion /Emilia-Datasetgated Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.audiotext-to-speech10M<n<100M489 likes48k downloads2y agoHugging Face07MLCommons /peoples_speech Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.audioautomatic-speech-recognition1M<n<10M285 likes39k downloads2y agoHugging Face08disco-eth /EuroSpeech EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper. This dataset is 16 kHz, the 24 kHz version of EuroSpeech can be found at… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech.audioautomatic-speech-recognition10M<n<100M100 likes36k downloads5mo agoHugging Face09facebook /multilingual_librispeech Dataset Card for MultiLingual LibriSpeech Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.audioautomatic-speech-recognition1M<n<10M190 likes36k downloads2y agoHugging Face10sarulab-speech /yodas2_sidon YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.audiotext-to-speech1M<n<10M65 likes32k downloads10mo agoHugging Face11twangodev /librivox-mirror LibriVox Mirror Fast, structured, continuously updated LibriVox audio mirror. Current snapshot Metric Value Published books 21,725 Published sections 493,206 Audio hours 132,555.1 Audio languages 86 Quarantined books 609 Last updated (UTC) 2026-09-22T13:24:09.765701Z Audio by language Language Hours English 131,605.7 German 417.0 Spanish 160.9 French 103.8 Portuguese 37.4 Polish 34.1 Dutch 25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.audioautomatic-speech-recognition100K<n<1M0 likes23k downloads3h agoHugging Face12ARTPARK-IISc /VaanigatedVAANI is an India-representative multi-modal multi-lingual dataset. The current version (phase 1- 80 districts, phase 2- 85 districts) contains ~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts, talking about 288K images covering 105 languages. From this audio data, 2,122 hours of transcribed data(text) is available, spanning almost evenly across the 165 districts. Project Vaani, by IISc, Bangalore and ARTPARK, is capturing the true diversity of India’s… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani.audioautomatic-speech-recognition1M<n<10M155 likes19k downloads5d agoHugging Face13simon3000 /genshin-voice Genshin Voice Genshin Voice is a dataset of voice lines from the popular game Genshin Impact. Hugging Face 🤗 Genshin-Voice ModelScope Genshin-Voice Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index. Last update at 2026-08-13 654252 wavs 7291 without speaker (1%) 52693 without transcription (8%) 1088 without inGameFilename (0%) Dataset Details Dataset Description The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.audioaudio-classification100K<n<1M271 likes18k downloads22d agoHugging Face14speechcolab /gigaspeechgated Dataset Card for Gigaspeech Dataset Description GigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training. The transcribed audio data is collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc. Example Usage The training split has several configurations of… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech.audioautomatic-speech-recognition10M<n<100M173 likes18k downloads8mo agoHugging Face15Sinoosoida /SpeechRu Russian Podcasts (unlabeled) ~186k unlabeled Russian-language podcast episodes scraped from the web, packaged as Parquet shards with the audio bytes embedded. The audio has no transcripts — this is an unsupervised / self-supervised audio corpus, suitable for ASR pre-training, speech-representation learning, TTS data mining, audio classification, and similar tasks. Each row contains: audio — the podcast episode (MP3, mostly 128 kbps / 44.1 kHz stereo), decoded on-the-fly via the… See the full description on the dataset page: https://huggingface.co/datasets/Sinoosoida/SpeechRu.audioautomatic-speech-recognition100K<n<1M4 likes17k downloads2mo agoHugging Face16PolyAI /minds14 MInDS-14 MINDS-14 is training and evaluation resource for intent detection task with spoken data. It covers 14 intents extracted from a commercial system in the e-banking domain, associated with spoken examples in 14 diverse language varieties. Example MInDS-14 can be downloaded and used as follows: from datasets import load_dataset minds_14 = load_dataset("PolyAI/minds14", "fr-FR") # for French # to download all data for multi-lingual fine-tuning uncomment following… See the full description on the dataset page: https://huggingface.co/datasets/PolyAI/minds14.audioautomatic-speech-recognition10K<n<100K108 likes15k downloads1y agoHugging Face17google /WaxalNLP Waxal Datasets The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus. Dataset Description The Waxal project provides datasets for both Automated Speech Recognition (ASR) and Text-to-Speech (TTS) for African languages. The goal of this dataset's creation and release is to facilitate research that improves the accuracy and fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/google/WaxalNLP.audioautomatic-speech-recognition1M<n<10M286 likes14k downloads21d agoHugging Face18PedroDKE /LibriS2S LibriS2S This repo contains scripts and alignment data to create a dataset build further upon librivoxDeEn such that it contains (German audio, German transcription, English audio, English transcription) quadruplets and can be used for Speech-to-Speech translation research. Because of this, the alignments are released under the same Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License These alignments were collected by downloading the English audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/PedroDKE/LibriS2S.audiotext-to-speech10K<n<100K4 likes12k downloads1y agoHugging Face19NCSpeech /YO-CPT-ru YO-CPT-ru YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily quality-filtered corpus of Russian speech mined from YouTube (via YODAS2) and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.audiotext-to-speech1M<n<10M16 likes12k downloads2mo agoHugging Face20edinburghcstr /ami Dataset Card for AMI Dataset Description The AMI Meeting Corpus consists of 100 hours of meeting recordings. The recordings use a range of signals synchronized to a common timeline. These include close-talking and far-field microphones, individual and room-view video cameras, and output from a slide projector and an electronic whiteboard. During the meetings, the participants also have unsynchronized pens available to them that record what is written. The meetings were… See the full description on the dataset page: https://huggingface.co/datasets/edinburghcstr/ami.audioautomatic-speech-recognition100K<n<1M96 likes12k downloads8mo agoHugging Face21ElectronicHug /short_video_ocr_dataset Short Video OCR / ASR Dataset An actively curated research dataset for building OCR, ASR, subtitle-alignment, and video-transcript pipelines for short social videos. It combines source videos and extracted frames with human review artifacts and model-generated text candidates. The primary languages are Ukrainian and Russian; English or mixed-language content may also occur. Status: work in progress. Model outputs and pseudo-label candidates are not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.imageimage-to-text1K<n<10K0 likes11k downloads3h agoHugging Face22projectkaira /Pretraining-V1 Indic TTS Unified v1 A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This dataset consolidates 17 distinct source datasets into a single, schema-normalized resource covering Indian / South Asian languages, plus major European, African, MENA, and Central Asian languages, with over 13.7 million utterances and 26,000+ hours of audio. All audio is resampled to 24 kHz mono. Every row follows an identical schema regardless of source… See the full description on the dataset page: https://huggingface.co/datasets/projectkaira/Pretraining-V1.audiotext-to-speech10M<n<100M0 likes11k downloads2mo agoHugging Face23gauduc /ulatroi 🧠 Project SLOB: Spontaneous Lifestyle & Observational Behaviors Dataset 📌 Abstract Welcome to the primary data ingestion node for Project SLOB. This repository hosts a massive, high-fidelity multimodal dataset designed to train next-generation Artificial Intelligence in recognizing, analyzing, and predicting spontaneous human behaviors in unconstrained, real-world video streams. This node is strictly used for the Spatial-Temporal Audio-Visual Synchronization… See the full description on the dataset page: https://huggingface.co/datasets/gauduc/ulatroi.textvideo-classificationn<1K0 likes10k downloads3mo agoHugging Face24CoRal-project /coral-v3gated CoRal: Danish Conversational and Read-aloud Dataset Version 3.0 Dataset Overview CoRal is a comprehensive Automatic Speech Recognition (ASR) dataset designed to capture the diversity of the Danish language across various dialects, accents, genders, and age groups. The primary goal of the CoRal dataset is to provide a robust resource for training and evaluating ASR models that can understand and transcribe spoken Danish in all its variations. Key Features… See the full description on the dataset page: https://huggingface.co/datasets/CoRal-project/coral-v3.audioautomatic-speech-recognition100K<n<1M6 likes9.9k downloads7mo agoHugging Face25freds0 /TAGARELA TAGARELA: A Portuguese Speech Dataset From Podcasts TAGARELA is a large-scale Portuguese speech dataset built from podcast audio and curated for speech technology research, especially Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The dataset contains more than 8,972 hours of Portuguese speech derived from the Cem Mil Podcasts collection. It includes Brazilian Portuguese and European Portuguese speech, processed through a pipeline involving audio standardization… See the full description on the dataset page: https://huggingface.co/datasets/freds0/TAGARELA.audioautomatic-speech-recognition1M<n<10M10 likes8.7k downloads2mo agoHugging Face26Digital-Divide-Data /Luhya-ASR-Data-subset-642H Luhya ASR Data Subset 642H Luhya speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M1 likes8.4k downloads1mo agoHugging Face27Bretagne /Banque_Sonore_Dialectes_Bretons [!NOTE] Dataset origin: http://banque.sonore.breton.free.fr/ Description Issue du site Banque Sonore des Dialectes Bretons Présentation du projet La Banque Sonore des Dialectes Bretons est un projet expérimental qui réunit sur internet un vaste ensemble d'enregistrements d'enquêtes effectuées depuis plus d'une dizaine d'années auprès de locuteurs traditionnels de breton. Alimentées par une équipe de bénévoles partageant un intérêt commun pour… See the full description on the dataset page: https://huggingface.co/datasets/Bretagne/Banque_Sonore_Dialectes_Bretons.audioautomatic-speech-recognition1K<n<10K3 likes7.6k downloads2mo agoHugging Face28LEMAS-Project /LEMAS-Dataset-train Overview This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project). It contains a large-scale training set (150k+ hours) and a curated evaluation set (500 utterances per language) covering 10 languages, all with word-level alignment. Fields key: unique utterance identifier; the first two characters indicate the language ID audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.texttext-to-speech100M<n<1B89 likes7.4k downloads6mo agoHugging Face29alexandrainst /ftspeech Dataset Card for FT Speech Dataset Summary This dataset is an upload of the FT Speech dataset. The training, validation and test splits are the original ones. Supported Tasks and Leaderboards Training automatic speech recognition is the intended task for this dataset. No leaderboard is active at this point. Languages The dataset is available in Danish (da). Dataset Structure Data Instances Size of downloaded dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/ftspeech.audioautomatic-speech-recognition1M<n<10M9 likes7.1k downloads2y agoHugging Face30KBLab /rixvox-v2 RixVox-v2: A Swedish parliamentary speech dataset RixVox-v2 is a parliamentary speech dataset spanning nearly 23000 hours of speech. The dataset was built by matching and force aligning speeches in parliamentary protocols to media recordings of debates. Each observation contains metadata about the speaker's name, gender, district, role, party affiliation, and the date the speech was given. We include identifiers for protocols, speeches and speakers that allow linking observations in… See the full description on the dataset page: https://huggingface.co/datasets/KBLab/rixvox-v2.audioautomatic-speech-recognition1M<n<10M12 likes7k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.