CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yasalma /tat_youtubeaudiotext-to-speech100K<n<1M0 likes5.3k downloads1y agoHugging Face02mesolitica /pseudolabel-malaysian-youtube-whisper-large-v3 Pseudolabel Malaysian Youtube videos using Whisper Large V3 Original dataset at https://huggingface.co/datasets/malaysia-ai/crawl-youtube, distributed pseudolabelled using 4x A100s script at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text-semisupervised/pseudolabel-whisper Each audio is 30 seconds. Each audio saved in 16k sample rate. audioautomatic-speech-recognition3 likes1.2k downloads3y agoHugging Face03alvanlii /cantonese-youtube-ttsgated Cantonese Audio TTS Dataset This dataset contains alvanlii/cantonese-radio, alvanlii/cantonese-youtube, plus a dataset of equal size. It is catered towards TTS (text-to-speech) use cases, more than the 2 previously published datasets, as there is more extensive filtering and audio enhancement. For speaker labelling, you can use speaker embedding models like Nvidia's TitaNet Filtered out: Overlapped voices, detected using pyannote/speaker-diarization-3.1 Music, detected using a… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube-tts.audiotext-to-speech1M<n<10M3 likes873 downloads6mo agoHugging Face04Scicom-intl /YouTube-Cantonese-Emilia YouTube Cantonese — Emilia 2,064,679 speaker-homogeneous Cantonese speech segments — 5,312.6 hours — produced by running alvanlii/cantonese-youtube through the Emilia speech-data pipeline (source separation → diarization → VAD segmentation → ASR → MOS filtering). Each row is one clean, single-speaker segment of 3–30 s with a transcript, a speaker turn label and a DNSMOS quality score. Audio is shipped separately as MP3s inside zip parts, in both an original and a silence-trimmed… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/YouTube-Cantonese-Emilia.tabularautomatic-speech-recognition1M<n<10M1 likes731 downloads1mo agoHugging Face05Rijgersberg /YouTube-Commons-nl-audio YouTube Commons NL Audio This dataset has audio files for the Dutch-language videos in Rijgersberg/YouTube-Commons-nl-transcriptions, all under a CC BY 4.0 license. It contains 11,669 files for a total runtime of 2493h 43m 5s, coming in at approximately 130 GB. Source The original source of the dataset (minus the titles, descriptions and audio files) is YouTube Commons: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons-nl-audio.audioautomatic-speech-recognition10K<n<100K1 likes725 downloads3d agoHugging Face06alvanlii /cantonese-youtubegated Cantonese Youtube Pseudo-Transcription Dataset Contains approximately 10k hours of audio sourced from YouTube Videos are chosen at random, and scraped on a channel basis Includes news, vlogs, entertainment, stories, health Columns transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube.audioautomatic-speech-recognition1M<n<10M51 likes607 downloads2y agoHugging Face07PerSets /youtube-persian-asrThis dataset consists of over 385 hours of audio extracted from various YouTube videos in the Persian language. Note: This dataset contains raw, unvalidated transcriptions. Users are advised to: 1. Perform their own quality assessment 2. Create their own train/validation/test splits based on their specific needs 3. Validate a subset of the data if needed for their use caseautomatic-speech-recognition7 likes464 downloads2y agoHugging Face08openbank-uz /youtube_transcriptions Dataset Description A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google's Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings. Use Cases Automatic Speech Recognition (ASR) for Uzbek Text-to-Speech (TTS) synthesis for Uzbek Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS) Speaker-conditioned TTS… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/youtube_transcriptions.audioautomatic-speech-recognition100K<n<1M2 likes463 downloads6mo agoHugging Face09islomov /news_youtube_uzbek_speech_dataset News Youtube Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/islomov/news_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K10 likes450 downloads1y agoHugging Face10OrcinusOrca /YouTube-Cantonese Cantonese Audio Dataset from YouTube This dataset contains Cantonese audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models. Data Source and Processing The data was obtained through the following process: Download: Audio (.m4a) and available Cantonese subtitles (.srt for zh-TW, zh-HK, zh-Hant)… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-Cantonese.audioautomatic-speech-recognition100K<n<1M5 likes445 downloads1y agoHugging Face11anhtunguyen98 /vi-asr-youtube-1582h vi-asr-youtube-1693h Bo du lieu ASR tieng Viet cat tu audio YouTube. Nhan sinh boi MAI-Transcribe-2 (Azure Speech) voi timestamp muc tu, loc bang mot model Zipformer doc lap. So doan 694,936 Tong thoi luong 1,693.4 gio Dinh dang MP3 64 kbps, 16 kHz, mono Video nguon 7,683 Kenh 13 Loai cat So doan Gio Do dai TB Muc dich long 493,045 1,439.9 10.5 s Doc dai lien tuc short 201,891 253.6 4.5 s Dictation, cau ngan Cau truc… See the full description on the dataset page: https://huggingface.co/datasets/anhtunguyen98/vi-asr-youtube-1582h.automatic-speech-recognition100K<n<1M0 likes440 downloads16d agoHugging Face12islomov /it_youtube_uzbek_speech_dataset IT Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language and with some english to better generalization. The data was collected from publicly available videos on YouTube related to the Information Technology (IT) field. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Mohir Dev YouTube channel (respect to the team for… See the full description on the dataset page: https://huggingface.co/datasets/islomov/it_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K7 likes387 downloads1y agoHugging Face13islomov /podcasts_tashkent_dialect_youtube_uzbek_speech_dataset Tashkent dialect focused podcasts youtube uzbek speech Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/islomov/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K6 likes377 downloads1y agoHugging Face14pourmand1376 /asr-farsi-youtube-chunked-30-seconds How To Use from datasets import load_dataset train = load_dataset('pourmand1376/asr-farsi-youtube-chunked-30-seconds', split='train+val') test =load_dataset('pourmand1376/asr-farsi-youtube-chunked-30-seconds', split='test') +300 Hours ASR dataset generated from this kaggle dataset audioautomatic-speech-recognition10K<n<100K11 likes278 downloads3y agoHugging Face15MohammadGholizadeh /youtube-farsi 📚 Unified Persian YouTube ASR Dataset (msghol/youtube-farsi) This dataset is an enhanced and user-ready version of PerSets/youtube-persian-asr, restructured for seamless integration with Hugging Face Dataset Viewer and downstream ASR pipelines. It simplifies the data format by combining audio and transcription into unified records, removing the need for preprocessing scripts. 🔍 Overview The dataset provides Persian-language audio-transcription pairs sourced from… See the full description on the dataset page: https://huggingface.co/datasets/MohammadGholizadeh/youtube-farsi.audioautomatic-speech-recognition100K<n<1M7 likes208 downloads1y agoHugging Face16azimislom /news_youtube_uzbek_speech_dataset News Youtube Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/azimislom/news_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K0 likes177 downloads7mo agoHugging Face17BoburAmirov /news_youtube_uzbek_speech_dataset News Youtube Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/BoburAmirov/news_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K0 likes171 downloads10mo agoHugging Face18hostbot77 /news_youtube_uzbek_speech_dataset News Youtube Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/hostbot77/news_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K0 likes158 downloads6mo agoHugging Face19softcatala /catalan-youtube-speech Catalan YouTube Speech Corpus This dataset contains 231,684 short audio clips of spontaneous Catalan speech, automatically extracted from public YouTube videos. Each clip is paired with two independent machine-generated transcription candidates, along with speaker gender, clip timing, and the source video's reuse license. It was built and published by Softcatalà, the volunteer organization behind free/open-source Catalan-language software. Homepage: https://www.softcatala.org/… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/catalan-youtube-speech.automatic-speech-recognition100K<n<1M3 likes148 downloads2mo agoHugging Face20Infatoshi /phonon-youtube-technical phonon-youtube-technical This dataset contains Phonon-authored manifests, labels, term/context indexes, attribution, and local rebuild scripts for technical YouTube speech data. It intentionally contains no audio files. Contents manifests/: segment timing, source URL, video ID, YouTube-reported license, and audio SHA-256 references. labels/: Phonon-authored label queues and teacher metadata. term_context_indexes/: technical term/context indexes used for analysis… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/phonon-youtube-technical.automatic-speech-recognition0 likes112 downloads4mo agoHugging Face21BSC-LT /distilled-catalan-youtube-speechThe Distilled Catalan YouTube Speech Corpus is the result of the automatic validation of the original Catalan YouTube Speech Corpus created by SoftCatala and shared trhough this HF repo: https://huggingface.co/datasets/softcatala/catalan-youtube-speech The corpus is -distilled- because only recordings with a high certaintity to be correct were taken and the rest were rejected.automatic-speech-recognition100K<n<1M1 likes97 downloads5mo agoHugging Face22Anilosan15 /YouTube_Video_Transkriptleri_TR Dataset Summary This dataset consists of nearly 5 hours of video from over 40 Creative Commons-licensed videos on YouTube. The videos contain the voices of more than 100 different people. The audio files have been resampled to 16 kHz. The videos have been divided into chunks of up to 25 seconds. This dataset is intended for developing Turkish STT (Speech-to-Text) models. Datasets Preparetion The audio files and transcript data were scraped from YouTube. The scraped… See the full description on the dataset page: https://huggingface.co/datasets/Anilosan15/YouTube_Video_Transkriptleri_TR.audioautomatic-speech-recognitionn<1K2 likes89 downloads2y agoHugging Face23OrcinusOrca /YouTube-English English Audio Dataset from YouTube This dataset contains English audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models. Data Source and Processing The data was obtained through the following process: Download: Audio (.m4a) and available English subtitles (.srt for en, en.j3PyPqV-e1s) were… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-English.audioautomatic-speech-recognition100K<n<1M2 likes75 downloads1y agoHugging Face24phongdtd /youtube_casual_audio\automatic-speech-recognition4 likes72 downloads2y agoHugging Face25Peacockery /tajik-asr-youtube tajik-asr-youtube Every Tajik-language clip from this project's YouTube scrape — 41 channels of news, talk shows, podcasts, audiobooks, and learning content — with machine transcripts and the verification scores left in as columns instead of applied as a filter. Pick your own quality threshold; the training corpus this project actually ships (tajik-asr-corpus-v3) is the gated subset. Layout Parquet shards under data/ with an audio struct column (16 kHz mono FLAC… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-youtube.tabularautomatic-speech-recognition100K<n<1M0 likes59 downloads3mo agoHugging Face26awaaz-se-alfaaz /YouTube-Evaluation-Set Awaaz se Alfaaz — YouTube Evaluation Set This dataset is the realistic multi-speaker evaluation set used in Awaaz se Alfaaz, accepted at LaTeLL 2026 — "Enhancing Urdu ASR with Whisper v3: Fine-Tuning on Latest Datasets and Realistic Multi-Speaker Evaluation with SLM Post-Processing." It contains 30 short-form Urdu YouTube videos (YouTube Shorts) covering a mix of news, sports, and current affairs content, along with human annotated gold transcripts and transcripts produced by… See the full description on the dataset page: https://huggingface.co/datasets/awaaz-se-alfaaz/YouTube-Evaluation-Set.textautomatic-speech-recognitionn<1K0 likes57 downloads2mo agoHugging Face27Infatoshi /phonon-youtube-cc-audio phonon-youtube-cc-audio This dataset contains the Creative-Commons subset of the Phonon technical YouTube speech corpus with audio segments included. It includes only source videos reported by yt-dlp as Creative Commons licensed and carries per-video attribution and license fields. Contents audio/: 16 kHz mono WAV segment files. manifests/segments.jsonl: one row per segment with local audio path, source URL, video ID, segment timing, SHA-256, and per-video… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/phonon-youtube-cc-audio.audioautomatic-speech-recognition1K<n<10K0 likes55 downloads4mo agoHugging Face28IbrahimDayax /somali-asr-synthetic-youtube Somali ASR Synthetic YouTube Dataset A Somali-language speech dataset derived from YouTube audio, intended for training and evaluating automatic speech recognition (ASR) and speech-to-text (STT) models. Transcriptions were generated synthetically (silver-standard) via ASR bootstrapping. Dataset Summary Split Samples train ~4,393 validation 200 test 100 Total ~4,693 Language: Somali (so) Audio format: WAV, 16 kHz, mono, 16-bit PCM Total… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-asr-synthetic-youtube.audioautomatic-speech-recognition1K<n<10K1 likes50 downloads4mo agoHugging Face29BoburAmirov /it_youtube_uzbek_speech_dataset IT Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language and with some english to better generalization. The data was collected from publicly available videos on YouTube related to the Information Technology (IT) field. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Mohir Dev YouTube channel (respect to the team for… See the full description on the dataset page: https://huggingface.co/datasets/BoburAmirov/it_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K2 likes47 downloads10mo agoHugging Face30ketav /hindi-youtube-asr-transcripts Hindi YouTube ASR Transcripts Auto-generated YouTube transcripts (VTT) from 21 Hindi channels for training ASR and TTS models. Quick Start # Download and extract wget https://huggingface.co/datasets/ketav/hindi-youtube-asr-transcripts/resolve/main/youtube_asr_data.tar.gz tar -xzf youtube_asr_data.tar.gz Stats Metric Value Channels 21 Total videos 109,981 Total hours 22,186.3 Hindi subtitles 98,309 Usable hours 19,032.5 Period 2025-2026… See the full description on the dataset page: https://huggingface.co/datasets/ketav/hindi-youtube-asr-transcripts.automatic-speech-recognition10K<n<100K0 likes41 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.