CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01espnet /yodas-granary Dataset Card for YODAS-Granary Repository: NeMo-speech-data-processor: Granary Paper: Granary: Speech Recognition and Translation Dataset in 25 European Languages Shared by: ESPnet Dataset Description YODAS-Granary is a curated subset of the larger nvidia/Granary dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages. Overview… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas-granary.audioautomatic-speech-recognition10M<n<100M33 likes77k downloads1y agoHugging Face02sarulab-speech /yodas2_sidon YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.audiotext-to-speech1M<n<10M65 likes32k downloads10mo agoHugging Face03NCSpeech /YO-CPT-ru YO-CPT-ru YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily quality-filtered corpus of Russian speech mined from YouTube (via YODAS2) and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.audiotext-to-speech1M<n<10M16 likes11k downloads2mo agoHugging Face04yasalma /tat_youtubeaudiotext-to-speech100K<n<1M0 likes5.3k downloads1y agoHugging Face05TheAgenticDataCompany /open-yap-1k Open Yap 1K: 1,000 hours of full-duplex natural conversation, free for commercial use Today we're releasing Open Yap 1K: 1,000 hours of dual-channel English conversation, capturing how people speak together naturally in real-world environments recorded in 48kHz. The dataset ships free for both commercial and research use. The sample on the Hugging Face Hub - 8.9 hours, 16 conversations, CC-BY-4.0, listenable in the dataset viewer. The full corpus - 1,000 hours, 1,602… See the full description on the dataset page: https://huggingface.co/datasets/TheAgenticDataCompany/open-yap-1k.audioaudio-to-audion<1K105 likes5.1k downloads19d agoHugging Face06espnet /yodas_owsmv4🏆 News: Our OWSM v4 paper won the Best Student Paper Award at INTERSPEECH 2025! Dataset Card for YODAS_OWSMv4 Paper: OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning (Best Student Paper at INTERSPEECH 2025) Authors: Yifan Peng, Muhammad Shakeel, Yui Sudo, William Chen, Jinchuan Tian, Chyi-Jiunn Lin, Shinji Watanabe Data Cleaning Scripts: ESPnet Model Demo: Gradio Dataset Description Open Whisper-style Speech Model (OWSM)is the first… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas_owsmv4.imageautomatic-speech-recognitionn<1K18 likes3.3k downloads1y agoHugging Face07TTS-AGI /emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset. https://huggingface.co/datasets/amphion/Emilia-Dataset audiotext-to-speech10M<n<100M5 likes3.1k downloads2y agoHugging Face08overflowwwww /yt-danish-public-v2audioaudio-classification100K<n<1M0 likes2.7k downloads2y agoHugging Face09retkowski /ytseg YTSeg: A Benchmark for Audio Chaptering and Video Transcript Segmentation We present YTSeg, a topically and structurally diverse benchmark for the audio chaptering and transcript segmentation task based on YouTube videos. The dataset comprises 19,299 videos from 393 channels, amounting to 6,533 content hours. The topics are wide-ranging, covering domains such as science, lifestyle, politics, health, economy, and technology. The videos are from various types of content formats… See the full description on the dataset page: https://huggingface.co/datasets/retkowski/ytseg.audiotoken-classification100K<n<1M8 likes2.7k downloads2mo agoHugging Face10yihao005 /Multi-Talker-SD Dataset Card for Multi-Talker-SD Dataset Description Multi-Talker-SD is a large-scale bilingual (English–Mandarin) multi-speaker meeting dataset designed to support research on speaker diarization and meeting transcription. Size: 1,000 simulated meetings Participants per meeting: 10–30 speakers Average duration: ~20 minutes per meeting, up to one hour Languages: English, Mandarin (code-switching possible) Audio characteristics: realistic speaker overlap… See the full description on the dataset page: https://huggingface.co/datasets/yihao005/Multi-Talker-SD.audioautomatic-speech-recognition4 likes1.4k downloads1y agoHugging Face11NCSpeech /YO-CPT-kk YO-CPT-kk YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the voice and, where… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-kk.audiotext-to-speech100K<n<1M10 likes1.3k downloads2mo agoHugging Face12mesolitica /pseudolabel-malaysian-youtube-whisper-large-v3 Pseudolabel Malaysian Youtube videos using Whisper Large V3 Original dataset at https://huggingface.co/datasets/malaysia-ai/crawl-youtube, distributed pseudolabelled using 4x A100s script at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text-semisupervised/pseudolabel-whisper Each audio is 30 seconds. Each audio saved in 16k sample rate. audioautomatic-speech-recognition3 likes1.3k downloads3y agoHugging Face13yasalma /audiobooks170 hours of aligned audiobooks taken from tatkniga.ru. There are 4 speakers with 17+ hours of audio and 20 speakers in total. All the books are in free access and most of them in public domain. audiotext-to-speech1K<n<10K0 likes845 downloads1y agoHugging Face14alvanlii /cantonese-youtube-ttsgated Cantonese Audio TTS Dataset This dataset contains alvanlii/cantonese-radio, alvanlii/cantonese-youtube, plus a dataset of equal size. It is catered towards TTS (text-to-speech) use cases, more than the 2 previously published datasets, as there is more extensive filtering and audio enhancement. For speaker labelling, you can use speaker embedding models like Nvidia's TitaNet Filtered out: Overlapped voices, detected using pyannote/speaker-diarization-3.1 Music, detected using a… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube-tts.audiotext-to-speech1M<n<10M3 likes804 downloads6mo agoHugging Face15Scicom-intl /YouTube-Cantonese-Emilia YouTube Cantonese — Emilia 2,064,679 speaker-homogeneous Cantonese speech segments — 5,312.6 hours — produced by running alvanlii/cantonese-youtube through the Emilia speech-data pipeline (source separation → diarization → VAD segmentation → ASR → MOS filtering). Each row is one clean, single-speaker segment of 3–30 s with a transcript, a speaker turn label and a DNSMOS quality score. Audio is shipped separately as MP3s inside zip parts, in both an original and a silence-trimmed… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/YouTube-Cantonese-Emilia.tabularautomatic-speech-recognition1M<n<10M1 likes790 downloads1mo agoHugging Face16yushin-ito /yodas-ja000 YODAS Japanese (ja000) Japanese manual caption subset of the YODAS dataset, repackaged for easier use. Source Original dataset: espnet/yodas (ja000 config) Paper: YODAS: YouTube-Oriented Dataset for Audio and Speech License: CC BY 3.0 Citation If you use this dataset, please cite the original YODAS paper: audioautomatic-speech-recognition100K<n<1M0 likes760 downloads6mo agoHugging Face17Yehor /audiobooks-xxlaudioautomatic-speech-recognition10M<n<100M1 likes672 downloads11mo agoHugging Face18yohannabelay /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/yohannabelay/fleurs.audioautomatic-speech-recognition100K<n<1M0 likes669 downloads2mo agoHugging Face19Rijgersberg /YouTube-Commons-nl-audio YouTube Commons NL Audio This dataset has audio files for the Dutch-language videos in Rijgersberg/YouTube-Commons-nl-transcriptions, all under a CC BY 4.0 license. It contains 11,669 files for a total runtime of 2493h 43m 5s, coming in at approximately 130 GB. Source The original source of the dataset (minus the titles, descriptions and audio files) is YouTube Commons: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons-nl-audio.audioautomatic-speech-recognition10K<n<100K1 likes666 downloads2d agoHugging Face20youvoi /WaxalNLP Waxal Datasets The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus. Dataset Description The Waxal project provides datasets for both Automated Speech Recognition (ASR) and Text-to-Speech (TTS) for African languages. The goal of this dataset's creation and release is to facilitate research that improves the accuracy and fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/youvoi/WaxalNLP.audioautomatic-speech-recognition1M<n<10M0 likes630 downloads3mo agoHugging Face21alvanlii /cantonese-youtubegated Cantonese Youtube Pseudo-Transcription Dataset Contains approximately 10k hours of audio sourced from YouTube Videos are chosen at random, and scraped on a channel basis Includes news, vlogs, entertainment, stories, health Columns transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube.audioautomatic-speech-recognition1M<n<10M51 likes604 downloads2y agoHugging Face22islomov /news_youtube_uzbek_speech_dataset News Youtube Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/islomov/news_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K10 likes472 downloads1y agoHugging Face23PerSets /youtube-persian-asrThis dataset consists of over 385 hours of audio extracted from various YouTube videos in the Persian language. Note: This dataset contains raw, unvalidated transcriptions. Users are advised to: 1. Perform their own quality assessment 2. Create their own train/validation/test splits based on their specific needs 3. Validate a subset of the data if needed for their use caseautomatic-speech-recognition7 likes459 downloads2y agoHugging Face24yujie-ovo /ChildTalk ChildTalk: A Multi-Dialect Chinese Child Speech Corpus with Full-Length Child–Caregiver Conversations for Speech Recognition 📖 Overview ChildTalk is a large-scale, publicly available multi-dialect Chinese child speech dialogue dataset. This dataset solves key problems found in existing Chinese child ASR corpora — mainly their small size, lack of natural conversations, limited dialectal coverage, and missing full-length dialogue recordings. It provides a solid… See the full description on the dataset page: https://huggingface.co/datasets/yujie-ovo/ChildTalk.automatic-speech-recognition2 likes458 downloads3mo agoHugging Face25michsethowusu /yoruba-speech-text-parallel Yoruba Speech-Text Parallel Dataset Dataset Description This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Yoruba - yo Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/yoruba-speech-text-parallel.audioautomatic-speech-recognition1M<n<10M3 likes447 downloads1y agoHugging Face26openbank-uz /youtube_transcriptions Dataset Description A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google's Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings. Use Cases Automatic Speech Recognition (ASR) for Uzbek Text-to-Speech (TTS) synthesis for Uzbek Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS) Speaker-conditioned TTS… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/youtube_transcriptions.audioautomatic-speech-recognition100K<n<1M2 likes443 downloads6mo agoHugging Face27espnet /yodas3 YODAS v3 Paper YODAS v3 is a large web-crawled dataset containing over 1.1 million hours of audio that were originally released under a CC-BY-3.0 license. The dataset contains audio in over 100 languages. YODAS v3 can be used for a variety of multi-modal tasks, including Automatic Speech Recognition, Text-to-Speech, and Audio Representation Learning. We crawl a distinct set of videos from the v1 and v2 versions of YODAS, to guarantee that there are no overlaps in the data. For… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas3.tabularaudio-to-audio1M<n<10M2 likes441 downloads6h agoHugging Face28OrcinusOrca /YouTube-Cantonese Cantonese Audio Dataset from YouTube This dataset contains Cantonese audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models. Data Source and Processing The data was obtained through the following process: Download: Audio (.m4a) and available Cantonese subtitles (.srt for zh-TW, zh-HK, zh-Hant)… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-Cantonese.audioautomatic-speech-recognition100K<n<1M5 likes439 downloads1y agoHugging Face29anhtunguyen98 /vi-asr-youtube-1582h vi-asr-youtube-1693h Bo du lieu ASR tieng Viet cat tu audio YouTube. Nhan sinh boi MAI-Transcribe-2 (Azure Speech) voi timestamp muc tu, loc bang mot model Zipformer doc lap. So doan 694,936 Tong thoi luong 1,693.4 gio Dinh dang MP3 64 kbps, 16 kHz, mono Video nguon 7,683 Kenh 13 Loai cat So doan Gio Do dai TB Muc dich long 493,045 1,439.9 10.5 s Doc dai lien tuc short 201,891 253.6 4.5 s Dictation, cau ngan Cau truc… See the full description on the dataset page: https://huggingface.co/datasets/anhtunguyen98/vi-asr-youtube-1582h.automatic-speech-recognition100K<n<1M0 likes438 downloads16d agoHugging Face30yongchanskii /OCW MIT OpenCourseWare dataset MIT OpenCourseWare dataset which consists of speech and its corresponding transcript. textautomatic-speech-recognition10K<n<100K0 likes432 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.