CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01espnet /yodas-granary Dataset Card for YODAS-Granary Repository: NeMo-speech-data-processor: Granary Paper: Granary: Speech Recognition and Translation Dataset in 25 European Languages Shared by: ESPnet Dataset Description YODAS-Granary is a curated subset of the larger nvidia/Granary dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages. Overview… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas-granary.audioautomatic-speech-recognition10M<n<100M33 likes77k downloads1y agoHugging Face02sarulab-speech /yodas2_sidon YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.audiotext-to-speech1M<n<10M65 likes32k downloads10mo agoHugging Face03NCSpeech /YO-CPT-ru YO-CPT-ru YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily quality-filtered corpus of Russian speech mined from YouTube (via YODAS2) and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.audiotext-to-speech1M<n<10M16 likes11k downloads2mo agoHugging Face04yasalma /tat_youtubeaudiotext-to-speech100K<n<1M0 likes5.3k downloads1y agoHugging Face05TheAgenticDataCompany /open-yap-1k Open Yap 1K: 1,000 hours of full-duplex natural conversation, free for commercial use Today we're releasing Open Yap 1K: 1,000 hours of dual-channel English conversation, capturing how people speak together naturally in real-world environments recorded in 48kHz. The dataset ships free for both commercial and research use. The sample on the Hugging Face Hub - 8.9 hours, 16 conversations, CC-BY-4.0, listenable in the dataset viewer. The full corpus - 1,000 hours, 1,602… See the full description on the dataset page: https://huggingface.co/datasets/TheAgenticDataCompany/open-yap-1k.audioaudio-to-audion<1K105 likes5.1k downloads19d agoHugging Face06TTS-AGI /emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset. https://huggingface.co/datasets/amphion/Emilia-Dataset audiotext-to-speech10M<n<100M5 likes3.1k downloads2y agoHugging Face07overflowwwww /yt-danish-public-v2audioaudio-classification100K<n<1M0 likes2.7k downloads2y agoHugging Face08retkowski /ytseg YTSeg: A Benchmark for Audio Chaptering and Video Transcript Segmentation We present YTSeg, a topically and structurally diverse benchmark for the audio chaptering and transcript segmentation task based on YouTube videos. The dataset comprises 19,299 videos from 393 channels, amounting to 6,533 content hours. The topics are wide-ranging, covering domains such as science, lifestyle, politics, health, economy, and technology. The videos are from various types of content formats… See the full description on the dataset page: https://huggingface.co/datasets/retkowski/ytseg.audiotoken-classification100K<n<1M8 likes2.7k downloads2mo agoHugging Face09yihao005 /Multi-Talker-SD Dataset Card for Multi-Talker-SD Dataset Description Multi-Talker-SD is a large-scale bilingual (English–Mandarin) multi-speaker meeting dataset designed to support research on speaker diarization and meeting transcription. Size: 1,000 simulated meetings Participants per meeting: 10–30 speakers Average duration: ~20 minutes per meeting, up to one hour Languages: English, Mandarin (code-switching possible) Audio characteristics: realistic speaker overlap… See the full description on the dataset page: https://huggingface.co/datasets/yihao005/Multi-Talker-SD.audioautomatic-speech-recognition4 likes1.4k downloads1y agoHugging Face10NCSpeech /YO-CPT-kk YO-CPT-kk YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the voice and, where… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-kk.audiotext-to-speech100K<n<1M10 likes1.3k downloads2mo agoHugging Face11mesolitica /pseudolabel-malaysian-youtube-whisper-large-v3 Pseudolabel Malaysian Youtube videos using Whisper Large V3 Original dataset at https://huggingface.co/datasets/malaysia-ai/crawl-youtube, distributed pseudolabelled using 4x A100s script at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text-semisupervised/pseudolabel-whisper Each audio is 30 seconds. Each audio saved in 16k sample rate. audioautomatic-speech-recognition3 likes1.3k downloads3y agoHugging Face12yasalma /audiobooks170 hours of aligned audiobooks taken from tatkniga.ru. There are 4 speakers with 17+ hours of audio and 20 speakers in total. All the books are in free access and most of them in public domain. audiotext-to-speech1K<n<10K0 likes845 downloads1y agoHugging Face13alvanlii /cantonese-youtube-ttsgated Cantonese Audio TTS Dataset This dataset contains alvanlii/cantonese-radio, alvanlii/cantonese-youtube, plus a dataset of equal size. It is catered towards TTS (text-to-speech) use cases, more than the 2 previously published datasets, as there is more extensive filtering and audio enhancement. For speaker labelling, you can use speaker embedding models like Nvidia's TitaNet Filtered out: Overlapped voices, detected using pyannote/speaker-diarization-3.1 Music, detected using a… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube-tts.audiotext-to-speech1M<n<10M3 likes804 downloads6mo agoHugging Face14yushin-ito /yodas-ja000 YODAS Japanese (ja000) Japanese manual caption subset of the YODAS dataset, repackaged for easier use. Source Original dataset: espnet/yodas (ja000 config) Paper: YODAS: YouTube-Oriented Dataset for Audio and Speech License: CC BY 3.0 Citation If you use this dataset, please cite the original YODAS paper: audioautomatic-speech-recognition100K<n<1M0 likes760 downloads6mo agoHugging Face15Yehor /audiobooks-xxlaudioautomatic-speech-recognition10M<n<100M1 likes672 downloads11mo agoHugging Face16yohannabelay /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/yohannabelay/fleurs.audioautomatic-speech-recognition100K<n<1M0 likes669 downloads2mo agoHugging Face17Rijgersberg /YouTube-Commons-nl-audio YouTube Commons NL Audio This dataset has audio files for the Dutch-language videos in Rijgersberg/YouTube-Commons-nl-transcriptions, all under a CC BY 4.0 license. It contains 11,669 files for a total runtime of 2493h 43m 5s, coming in at approximately 130 GB. Source The original source of the dataset (minus the titles, descriptions and audio files) is YouTube Commons: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons-nl-audio.audioautomatic-speech-recognition10K<n<100K1 likes666 downloads2d agoHugging Face18youvoi /WaxalNLP Waxal Datasets The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus. Dataset Description The Waxal project provides datasets for both Automated Speech Recognition (ASR) and Text-to-Speech (TTS) for African languages. The goal of this dataset's creation and release is to facilitate research that improves the accuracy and fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/youvoi/WaxalNLP.audioautomatic-speech-recognition1M<n<10M0 likes630 downloads3mo agoHugging Face19alvanlii /cantonese-youtubegated Cantonese Youtube Pseudo-Transcription Dataset Contains approximately 10k hours of audio sourced from YouTube Videos are chosen at random, and scraped on a channel basis Includes news, vlogs, entertainment, stories, health Columns transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube.audioautomatic-speech-recognition1M<n<10M51 likes604 downloads2y agoHugging Face20islomov /news_youtube_uzbek_speech_dataset News Youtube Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/islomov/news_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K10 likes472 downloads1y agoHugging Face21michsethowusu /yoruba-speech-text-parallel Yoruba Speech-Text Parallel Dataset Dataset Description This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Yoruba - yo Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/yoruba-speech-text-parallel.audioautomatic-speech-recognition1M<n<10M3 likes447 downloads1y agoHugging Face22openbank-uz /youtube_transcriptions Dataset Description A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google's Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings. Use Cases Automatic Speech Recognition (ASR) for Uzbek Text-to-Speech (TTS) synthesis for Uzbek Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS) Speaker-conditioned TTS… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/youtube_transcriptions.audioautomatic-speech-recognition100K<n<1M2 likes443 downloads6mo agoHugging Face23OrcinusOrca /YouTube-Cantonese Cantonese Audio Dataset from YouTube This dataset contains Cantonese audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models. Data Source and Processing The data was obtained through the following process: Download: Audio (.m4a) and available Cantonese subtitles (.srt for zh-TW, zh-HK, zh-Hant)… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-Cantonese.audioautomatic-speech-recognition100K<n<1M5 likes439 downloads1y agoHugging Face24speech-uk /yodas2 YODAS2 for 🇺🇦 Ukrainian Ukrainian validated subset of YODAS2 Community Discord: https://bit.ly/discord-uds Speech Recognition: https://t.me/speech_recognition_uk Speech Synthesis: https://t.me/speech_synthesis_uk Stats Total files processed: 400213 Total duration: 998h 41m 3s textautomatic-speech-recognition100K<n<1M2 likes427 downloads11mo agoHugging Face25Chalermdej /yodas2_sidon_th_tts Thai TTS Dataset — Filtered & Quality-Verified from YODAS2 sidon A filtered, quality-verified Thai text-to-speech dataset derived from sarulab-speech/yodas2_sidon, with transcriptions verified by multiple ASR models and Gemini, text fully normalized to Thai, and audio quality-screened with DNSMOS. Dataset Summary Samples 141,927 Audio hours 156.0 Speakers 4,199 Sample rate 24,000 Hz Format WAV, PCM 16-bit, mono Language Thai Source… See the full description on the dataset page: https://huggingface.co/datasets/Chalermdej/yodas2_sidon_th_tts.audiotext-to-speech100K<n<1M3 likes425 downloads4mo agoHugging Face26ymoslem /EUbookshop-Speech-Irish Dataset Details Synthetic audio dataset, created using Azure text-to-speech service. The bilingual text is a portion of the EUbookshop dataset, consisting of 33,634 text segments. The dataset includes two sets of audio data, one with a female voice (OrlaNeural) and the other with a male voice (ColmNeural). The speech data comprises approximately 159 hours and 45 minutes (159:45:05) spread across 67,268 utterances. Dataset Structure Dataset({ features: ['audio'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/EUbookshop-Speech-Irish.audioautomatic-speech-recognition10K<n<100K0 likes417 downloads2y agoHugging Face27yuekai /speechio_test SpeechIO ASR Test Sets (parquet) Parquet repackaging of the SpeechColab SpeechIO Mandarin ASR benchmark, re-exported from yuekai/speechio (Lhotse cuts) into standard HuggingFace parquet with embedded 16 kHz audio. 27 test sets: SPEECHIO_ASR_ZH00000 ... SPEECHIO_ASR_ZH00026, each a config with a single test split. ~43k utterances, ~66 hours total, evaluation only. Columns column type note segment_id string utterance id speaker string speaker id… See the full description on the dataset page: https://huggingface.co/datasets/yuekai/speechio_test.audioautomatic-speech-recognition10K<n<100K0 likes407 downloads3mo agoHugging Face28islomov /it_youtube_uzbek_speech_dataset IT Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language and with some english to better generalization. The data was collected from publicly available videos on YouTube related to the Information Technology (IT) field. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Mohir Dev YouTube channel (respect to the team for… See the full description on the dataset page: https://huggingface.co/datasets/islomov/it_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K7 likes401 downloads1y agoHugging Face29islomov /podcasts_tashkent_dialect_youtube_uzbek_speech_dataset Tashkent dialect focused podcasts youtube uzbek speech Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/islomov/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K6 likes396 downloads1y agoHugging Face30skypro1111 /whisper-dataset-ytb-uk Dataset Card for Dataset Name This dataset is collected from youtube. audioautomatic-speech-recognition10K<n<100K2 likes366 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.