CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yasalma /tat_youtubeaudiotext-to-speech100K<n<1M0 likes5.3k downloads1y agoHugging Face02alvanlii /cantonese-youtube-ttsgated Cantonese Audio TTS Dataset This dataset contains alvanlii/cantonese-radio, alvanlii/cantonese-youtube, plus a dataset of equal size. It is catered towards TTS (text-to-speech) use cases, more than the 2 previously published datasets, as there is more extensive filtering and audio enhancement. For speaker labelling, you can use speaker embedding models like Nvidia's TitaNet Filtered out: Overlapped voices, detected using pyannote/speaker-diarization-3.1 Music, detected using a… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube-tts.audiotext-to-speech1M<n<10M3 likes873 downloads6mo agoHugging Face03Scicom-intl /YouTube-Cantonese-Emilia YouTube Cantonese — Emilia 2,064,679 speaker-homogeneous Cantonese speech segments — 5,312.6 hours — produced by running alvanlii/cantonese-youtube through the Emilia speech-data pipeline (source separation → diarization → VAD segmentation → ASR → MOS filtering). Each row is one clean, single-speaker segment of 3–30 s with a transcript, a speaker turn label and a DNSMOS quality score. Audio is shipped separately as MP3s inside zip parts, in both an original and a silence-trimmed… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/YouTube-Cantonese-Emilia.tabularautomatic-speech-recognition1M<n<10M1 likes731 downloads1mo agoHugging Face04Rijgersberg /YouTube-Commons-nl-audio YouTube Commons NL Audio This dataset has audio files for the Dutch-language videos in Rijgersberg/YouTube-Commons-nl-transcriptions, all under a CC BY 4.0 license. It contains 11,669 files for a total runtime of 2493h 43m 5s, coming in at approximately 130 GB. Source The original source of the dataset (minus the titles, descriptions and audio files) is YouTube Commons: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons-nl-audio.audioautomatic-speech-recognition10K<n<100K1 likes725 downloads3d agoHugging Face05alvanlii /cantonese-youtubegated Cantonese Youtube Pseudo-Transcription Dataset Contains approximately 10k hours of audio sourced from YouTube Videos are chosen at random, and scraped on a channel basis Includes news, vlogs, entertainment, stories, health Columns transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube.audioautomatic-speech-recognition1M<n<10M51 likes607 downloads2y agoHugging Face06openbank-uz /youtube_transcriptions Dataset Description A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google's Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings. Use Cases Automatic Speech Recognition (ASR) for Uzbek Text-to-Speech (TTS) synthesis for Uzbek Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS) Speaker-conditioned TTS… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/youtube_transcriptions.audioautomatic-speech-recognition100K<n<1M2 likes463 downloads6mo agoHugging Face07islomov /news_youtube_uzbek_speech_dataset News Youtube Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/islomov/news_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K10 likes450 downloads1y agoHugging Face08OrcinusOrca /YouTube-Cantonese Cantonese Audio Dataset from YouTube This dataset contains Cantonese audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models. Data Source and Processing The data was obtained through the following process: Download: Audio (.m4a) and available Cantonese subtitles (.srt for zh-TW, zh-HK, zh-Hant)… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-Cantonese.audioautomatic-speech-recognition100K<n<1M5 likes445 downloads1y agoHugging Face09islomov /it_youtube_uzbek_speech_dataset IT Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language and with some english to better generalization. The data was collected from publicly available videos on YouTube related to the Information Technology (IT) field. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Mohir Dev YouTube channel (respect to the team for… See the full description on the dataset page: https://huggingface.co/datasets/islomov/it_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K7 likes387 downloads1y agoHugging Face10islomov /podcasts_tashkent_dialect_youtube_uzbek_speech_dataset Tashkent dialect focused podcasts youtube uzbek speech Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/islomov/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K6 likes377 downloads1y agoHugging Face11pourmand1376 /asr-farsi-youtube-chunked-30-seconds How To Use from datasets import load_dataset train = load_dataset('pourmand1376/asr-farsi-youtube-chunked-30-seconds', split='train+val') test =load_dataset('pourmand1376/asr-farsi-youtube-chunked-30-seconds', split='test') +300 Hours ASR dataset generated from this kaggle dataset audioautomatic-speech-recognition10K<n<100K11 likes278 downloads3y agoHugging Face12MohammadGholizadeh /youtube-farsi 📚 Unified Persian YouTube ASR Dataset (msghol/youtube-farsi) This dataset is an enhanced and user-ready version of PerSets/youtube-persian-asr, restructured for seamless integration with Hugging Face Dataset Viewer and downstream ASR pipelines. It simplifies the data format by combining audio and transcription into unified records, removing the need for preprocessing scripts. 🔍 Overview The dataset provides Persian-language audio-transcription pairs sourced from… See the full description on the dataset page: https://huggingface.co/datasets/MohammadGholizadeh/youtube-farsi.audioautomatic-speech-recognition100K<n<1M7 likes208 downloads1y agoHugging Face13azimislom /news_youtube_uzbek_speech_dataset News Youtube Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/azimislom/news_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K0 likes177 downloads7mo agoHugging Face14BoburAmirov /news_youtube_uzbek_speech_dataset News Youtube Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/BoburAmirov/news_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K0 likes171 downloads10mo agoHugging Face15hostbot77 /news_youtube_uzbek_speech_dataset News Youtube Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/hostbot77/news_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K0 likes158 downloads6mo agoHugging Face16Anilosan15 /YouTube_Video_Transkriptleri_TR Dataset Summary This dataset consists of nearly 5 hours of video from over 40 Creative Commons-licensed videos on YouTube. The videos contain the voices of more than 100 different people. The audio files have been resampled to 16 kHz. The videos have been divided into chunks of up to 25 seconds. This dataset is intended for developing Turkish STT (Speech-to-Text) models. Datasets Preparetion The audio files and transcript data were scraped from YouTube. The scraped… See the full description on the dataset page: https://huggingface.co/datasets/Anilosan15/YouTube_Video_Transkriptleri_TR.audioautomatic-speech-recognitionn<1K2 likes89 downloads2y agoHugging Face17OrcinusOrca /YouTube-English English Audio Dataset from YouTube This dataset contains English audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models. Data Source and Processing The data was obtained through the following process: Download: Audio (.m4a) and available English subtitles (.srt for en, en.j3PyPqV-e1s) were… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-English.audioautomatic-speech-recognition100K<n<1M2 likes75 downloads1y agoHugging Face18Peacockery /tajik-asr-youtube tajik-asr-youtube Every Tajik-language clip from this project's YouTube scrape — 41 channels of news, talk shows, podcasts, audiobooks, and learning content — with machine transcripts and the verification scores left in as columns instead of applied as a filter. Pick your own quality threshold; the training corpus this project actually ships (tajik-asr-corpus-v3) is the gated subset. Layout Parquet shards under data/ with an audio struct column (16 kHz mono FLAC… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-youtube.tabularautomatic-speech-recognition100K<n<1M0 likes59 downloads3mo agoHugging Face19awaaz-se-alfaaz /YouTube-Evaluation-Set Awaaz se Alfaaz — YouTube Evaluation Set This dataset is the realistic multi-speaker evaluation set used in Awaaz se Alfaaz, accepted at LaTeLL 2026 — "Enhancing Urdu ASR with Whisper v3: Fine-Tuning on Latest Datasets and Realistic Multi-Speaker Evaluation with SLM Post-Processing." It contains 30 short-form Urdu YouTube videos (YouTube Shorts) covering a mix of news, sports, and current affairs content, along with human annotated gold transcripts and transcripts produced by… See the full description on the dataset page: https://huggingface.co/datasets/awaaz-se-alfaaz/YouTube-Evaluation-Set.textautomatic-speech-recognitionn<1K0 likes57 downloads2mo agoHugging Face20IbrahimDayax /somali-asr-synthetic-youtube Somali ASR Synthetic YouTube Dataset A Somali-language speech dataset derived from YouTube audio, intended for training and evaluating automatic speech recognition (ASR) and speech-to-text (STT) models. Transcriptions were generated synthetically (silver-standard) via ASR bootstrapping. Dataset Summary Split Samples train ~4,393 validation 200 test 100 Total ~4,693 Language: Somali (so) Audio format: WAV, 16 kHz, mono, 16-bit PCM Total… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-asr-synthetic-youtube.audioautomatic-speech-recognition1K<n<10K1 likes50 downloads4mo agoHugging Face21BoburAmirov /it_youtube_uzbek_speech_dataset IT Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language and with some english to better generalization. The data was collected from publicly available videos on YouTube related to the Information Technology (IT) field. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Mohir Dev YouTube channel (respect to the team for… See the full description on the dataset page: https://huggingface.co/datasets/BoburAmirov/it_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K2 likes47 downloads10mo agoHugging Face22ErfanRou /youtube-300h-movies-sonioxgated Persian Speech Corpus — full three-pool release 309.72 hours · 157,279 clips · 736 source videos · Soniox transcripts on every clip. This is the complete quality-gated output of the persian-expressive-corpus pipeline. It is organised into three mutually exclusive pools. Read the pool column before using a clip — they are not interchangeable. pool clips hours transcript emotion label QC status recommended use A 66,108 108.34 yes yes, 7-class passed all gates expressive… See the full description on the dataset page: https://huggingface.co/datasets/ErfanRou/youtube-300h-movies-soniox.tabularautomatic-speech-recognition100K<n<1M0 likes36 downloads13d agoHugging Face23azimislom /podcasts_tashkent_dialect_youtube_uzbek_speech_dataset Tashkent dialect focused podcasts youtube uzbek speech Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/azimislom/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K0 likes32 downloads7mo agoHugging Face24lilgoose777 /nepali-youtube-datasetgated Nepali Speech Dataset (YouTube-sourced) 585 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split. Splits train: 585 segments validation: 0 segments test: 0 segments Transcript columns — read this before training Each segment carries three transcript variants. They are NOT interchangeable: text_original — the YouTube caption text (if any) that overlapped this… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepali-youtube-dataset.audioautomatic-speech-recognitionn<1K0 likes31 downloads16d agoHugging Face25DrIAmed /darija-youtube-dataset Darija YouTube Dataset A dataset of Moroccan Arabic (Darija) speech scraped from YouTube, with transcriptions in Arabic script, Latin script (Arabizi), and English translations. Dataset Description This dataset contains sentence-level audio segments of Darija speech, paired with: Arabic transcription (modern Moroccan Arabic script) Latin transliteration (Arabizi format: 3=ع, 7=ح, 9=ق, etc.) English translation Columns Column Type Description audio… See the full description on the dataset page: https://huggingface.co/datasets/DrIAmed/darija-youtube-dataset.audioautomatic-speech-recognition1K<n<10K0 likes22 downloads7mo agoHugging Face26yasalma /youtube_ttTotal duration: 2.2 hours audioautomatic-speech-recognition1K<n<10K0 likes21 downloads1y agoHugging Face27azimislom /it_youtube_uzbek_speech_dataset IT Uzbek Speech Dataset Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language and with some english to better generalization. The data was collected from publicly available videos on YouTube related to the Information Technology (IT) field. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Mohir Dev YouTube channel (respect to the team for… See the full description on the dataset page: https://huggingface.co/datasets/azimislom/it_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K0 likes19 downloads7mo agoHugging Face28cillegio /az-asr-youtube-136hgated Labelling field value label_origin asr:google speech_register spontaneous channel wideband-16k provenance documented Google ASR output over YouTube audio. No human labels anywhere in it. label_origin distinguishes text that existed before the audio (script, exact by construction) from text written by a listener (human-transcript, high but edited) from machine output (asr:<vendor>, bounded by that vendor's error rate). The older label_type field is retained… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-youtube-136h.textautomatic-speech-recognition10K<n<100K0 likes18 downloads6d agoHugging Face29BoburAmirov /podcasts_tashkent_dialect_youtube_uzbek_speech_dataset Tashkent dialect focused podcasts youtube uzbek speech Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/BoburAmirov/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K0 likes12 downloads10mo agoHugging Face30Zarakun /youtube_ua_noisy_subtitles_test The list of all subsets in the dataset Each subset is generated splitting videos from given particular ukrainiam YouTube channel All subsets are in test split "opodcast" subset is from channel "О! ПОДКАСТ" "rozdympodcast" subset is from channel "Роздум | Подкаст" "test" subset is just a small subset of samples Loading a particular subset >>> data_files = {"train": "data/<your_subset>.parquet"} >>> data = load_dataset("Zarakun/youtube_ua_subtitles_test"… See the full description on the dataset page: https://huggingface.co/datasets/Zarakun/youtube_ua_noisy_subtitles_test.tabularautomatic-speech-recognition1K<n<10K0 likes11 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.