CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sarulab-speech /yodas2_sidon YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.audiotext-to-speech1M<n<10M65 likes32k downloads10mo agoHugging Face02sarulab-speech /mls_sidon MLS-Sidon Overview This dataset is a cleansed version of Multilingual LibriSpeech (MLS) with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. The dataset is provided in WebDataset format for efficient large-scale training. Source: Multilingual LibriSpeech Languages: English, German, French, Spanish, Italian, Polish, Dutch, Portuguese Format: WebDataset (.tar shards) License: CC-BY-4.0 Dataset Structure Each sample in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/mls_sidon.audiotext-to-speech10M<n<100M11 likes9.8k downloads1y agoHugging Face03speechcolab /gigaspeech2gated Dataset Card for GigaSpeech 2 Dataset Description GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese. Repository: https://github.com/SpeechColab/GigaSpeech2 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.audioautomatic-speech-recognition10M<n<100M71 likes5.8k downloads6mo agoHugging Face04sarulab-speech /commonvoice22_sidongated CV22-Sidon Overview This dataset hosts a release of Mozilla Common Voice 22 restored with the Sidon speech restoration model. Source: Mozilla Common Voice 22.0 Processing: Sidon denoising (sarulab-speech/sidon-v0.1) with 21 s chunks and 48 kHz reconstruction Format: WebDataset shards (.tar.gz) Manifest: paths.yaml enumerates every shard path for Hugging Face–style loading License: Original Common Voice license (CC0 1.0) Languages 137 language folders are… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon.audiotext-to-speech10M<n<100M30 likes1.8k downloads1y agoHugging Face05litagin /Galgame_Speech_SER_16kHz Dataset Card for Galgame_Speech_SER_16kHz [!IMPORTANT]The following rules (in the original repository) must be followed: 必须遵守GNU General Public License v3.0内的所有协议!附加:禁止商用,本数据集以及使用本数据集训练出来的任何模型都不得用于任何商业行为,如要用于商业用途,请找数据列表内的所有厂商授权(笑),因违反开源协议而出现的任何问题都与本人无关! 训练出来的模型必须开源,是否在README内引用本数据集由训练者自主决定,不做强制要求。 English: You must comply with all the terms of the GNU General Public License v3.0!Additional note: Commercial use is prohibited. This dataset and any model trained using this dataset… See the full description on the dataset page: https://huggingface.co/datasets/litagin/Galgame_Speech_SER_16kHz.audioautomatic-speech-recognition1M<n<10M17 likes973 downloads2y agoHugging Face06litagin /reazon-speech-v2-clonegated Reazon Speech v2 dataset mirror Original Dataset Source Hugging Face Dataset Page: reazon-research/reazonspeech Project Page: Reazon Research License This dataset is a mirror of the original Reazon Speech v2 dataset, but on 🤗 server (so may be faster). This dataset is licensed under the CDLA-Sharing-1.0. The original dataset comes with the following restriction: TO USE THIS DATASET, YOU MUST AGREE THAT YOU WILL USE THE DATASET SOLELY FOR THE PURPOSE OF… See the full description on the dataset page: https://huggingface.co/datasets/litagin/reazon-speech-v2-clone.audioautomatic-speech-recognition10K<n<100K12 likes851 downloads2y agoHugging Face07laion /unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1audio100M<n<1B4 likes671 downloads1y agoHugging Face08akuzdeuov /qwen3-tts-multilingual-emotional-speechaudio1M<n<10M0 likes505 downloads11d agoHugging Face09litagin /Galgame_Speech_ASR_16kHz Dataset Card for Galgame_Speech_ASR_16kHz [!IMPORTANT]The following rules (in the original repository) must be followed: 必须遵守GNU General Public License v3.0内的所有协议!附加:禁止商用,本数据集以及使用本数据集训练出来的任何模型都不得用于任何商业行为,如要用于商业用途,请找数据列表内的所有厂商授权(笑),因违反开源协议而出现的任何问题都与本人无关! 训练出来的模型必须开源,是否在README内引用本数据集由训练者自主决定,不做强制要求。 English: You must comply with all the terms of the GNU General Public License v3.0!Additional note: Commercial use is prohibited. This dataset and any model trained using this dataset… See the full description on the dataset page: https://huggingface.co/datasets/litagin/Galgame_Speech_ASR_16kHz.audioautomatic-speech-recognition1M<n<10M48 likes450 downloads2y agoHugging Face10sambal /speech_400kaudio100K<n<1M0 likes439 downloads2y agoHugging Face11sheng22213 /multi_round_speech_180kaudio1M<n<10M2 likes426 downloads1y agoHugging Face12recursal /reprocessed_singapore_national_speech_corpus Dataset Card for Reprocessed National Speech Corpus NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here. Dataset Details Dataset Description Dataset Description: The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.audiotext-generation1M<n<10M7 likes340 downloads2y agoHugging Face13krishnakalyan3 /emo_speech_filtered_v12 second filtered emotional speech in webdataset format https://huggingface.co/datasets/EQ4You/Emotional_Speech audio10K<n<100K0 likes181 downloads2y agoHugging Face14sheng22213 /speech_text-tts_audioaudio10K<n<100K0 likes136 downloads1y agoHugging Face15sambal /speech_dataaudio1M<n<10M0 likes127 downloads2y agoHugging Face16gijs /speech-utterances Speech Utterances Dataset Collection A collection of human non-speech vocal sound datasets in WebDataset format, useful for audio classification tasks involving vocal expressions, emotions, and non-verbal sounds. Subsets all (default) All datasets concatenated together (~75k samples total). NonSpeech7k Samples: 7,014 (train: 6,289, test: 725) Classes: Breathing, Coughing, Crying, Laughing, Screaming, Sneezing, Yawning Source: Zenodo License: CC… See the full description on the dataset page: https://huggingface.co/datasets/gijs/speech-utterances.audioaudio-classification100K<n<1M1 likes112 downloads8mo agoHugging Face17issai /Multilingual_Speech_Dataset Multilingual Speech Dataset Paper: A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English Repository: https://github.com/IS2AI/MultilingualASR Description: This repository provides the dataset used in the paper "A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English". The paper focuses on training a single end-to-end (E2E) ASR model for Kazakh, Russian, and English, comparing monolingual and multilingual approaches… See the full description on the dataset page: https://huggingface.co/datasets/issai/Multilingual_Speech_Dataset.audioautomatic-speech-recognition100K<n<1M3 likes108 downloads2y agoHugging Face18Aalto-Speech-Synthesis /stortinget_speech_corpus_v1.0 Dataset Card for Stortinget Speech Corpus V1.0 Overview This is the WebDataset version of the Stortinget Speech Corpus V1.0, originally created by the National Library of Norway. We re-organize it into WebDataset format for better usability. The Stortinget Speech Corpus (SSC) is a 5000+ hours speech dataset for weak supervision ASR created from audio andaligned proceedings text from Stortinget, the Norwegian Parliament. For more information, please refer to the original… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/stortinget_speech_corpus_v1.0.audioautomatic-speech-recognition100K<n<1M0 likes83 downloads5mo agoHugging Face19cagatayn /multi_accent_speech Multi-Accent English Speech Corpus (Augmented & Speaker-Disjoint) This dataset is a curated and augmented multi-accent English speech corpus designed for speech recognition, accent classification, and representation learning.It consolidates multiple open-source accent corpora, converts all audio to a unified format, applies targeted data augmentation, and exports in a tidy, Hugging Face–ready structure. ✨ Key Features Accents covered (12 total):american_english… See the full description on the dataset page: https://huggingface.co/datasets/cagatayn/multi_accent_speech.audio100K<n<1M2 likes78 downloads1y agoHugging Face20freococo /myanmar-english-accent-speech Myanmar English Accent Speech (PVTV & FOEIM) This dataset contains English speech by Myanmar speakers, collected from public videos published by PVTV and FOEIM — two media channels operating under the National Unity Government (NUG). The clips reflect a wide range of spoken English contexts: interviews, announcements, sermons, and educational content. The speakers vary in tone, pace, and emotion — but all share the characteristic sound of Burmese-accented English. This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-accent-speech.audioautomatic-speech-recognition1K<n<10K1 likes59 downloads1y agoHugging Face21Reord-AI /english-casual-speech-sample-south-african-accentgated English Casual Speech Sample (South African Accent) South African crowd-sourced participants respond to questions about their daily lives and activities. This dataset is a sample of a larger collection from the same data collection campaign. Changelog FEB 2026: initial share. ASR (Chirp3) transcripts. WER: 12% Specs Speakers: ~550 unique South African speakers Total duration: ~60 hours Files sample rate: 48kHz Actual sample rate: TBD Language: English (SA… See the full description on the dataset page: https://huggingface.co/datasets/Reord-AI/english-casual-speech-sample-south-african-accent.audio10K<n<100K0 likes47 downloads8mo agoHugging Face22kehanlu /Speech-IFEvalaudio1K<n<10K0 likes44 downloads2y agoHugging Face23tracygu /efficient-speech-codec Efficient Speech Codec Paper Code Demo Page audio100K<n<1M0 likes39 downloads2y agoHugging Face24Peacockery /mozilla-common-voice-spontaneous-speech-asr-shared-task Mozilla Common Voice Spontaneous Speech ASR Shared Task This repository combines the Mozilla Data Collective Common Voice spontaneous speech ASR shared-task train/dev and test archives in one place. Locales present across the combined train/dev and test packages: ady, aln, bas, bew, bxk, cgg, el-CY, hch, kbd, kcn, koo, led, lke, lth, meh, mmc, pne, qxp, ruc, rwm, sco, tob, top, ttj, ukv, ush. Split package Mozilla Data Collective dataset ID Hub archive Original MDC archive… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/mozilla-common-voice-spontaneous-speech-asr-shared-task.textautomatic-speech-recognition10K<n<100K0 likes38 downloads3mo agoHugging Face25novateur /speech_datatsetaudio10K<n<100K0 likes35 downloads2y agoHugging Face26Subuday /lj_speechtextn<1K0 likes28 downloads3y agoHugging Face27onlysainaa /common-voice-scripted-speech-24.0-mongoliantext10K<n<100K0 likes26 downloads8mo agoHugging Face28zuhri025 /speech-utterances Speech Utterances Dataset Collection A collection of human non-speech vocal sound datasets in WebDataset format, useful for audio classification tasks involving vocal expressions, emotions, and non-verbal sounds. Subsets all (default) All datasets concatenated together (~75k samples total). NonSpeech7k Samples: 7,014 (train: 6,289, test: 725) Classes: Breathing, Coughing, Crying, Laughing, Screaming, Sneezing, Yawning Source: Zenodo License: CC… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/speech-utterances.audioaudio-classification100K<n<1M0 likes23 downloads8mo agoHugging Face29jlking /speechdialogue_dataaudio10K<n<100K0 likes22 downloads1y agoHugging Face30sambal /speech_40kaudio100K<n<1M0 likes21 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.