CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Whispering-GPT /linustechtips-transcript-audio Dataset Card for "linustechtips" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips. Data Fields The dataset is composed by: id: Id of the youtube video. channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.audioautomatic-speech-recognitionn<1K4 likes2.3k downloads4y agoHugging Face02Whispering-GPT /lex-fridman-podcast-transcript-audio Dataset Card for "lexFridmanPodcast-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast. Data Fields The dataset is composed by: id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.audioautomatic-speech-recognitionn<1K0 likes2.1k downloads4y agoHugging Face03ARTPARK-IISc /Vaani-transcription-partgatedThis dataset is part of the Vaani dataset and consists of only transcribed speech data. It has a total duration of 2041.54 hours, covering 59 languages. This table represents the audio and transcription duration data for various languages. Language Angami Angika Ao Assamese Awadhi Bajjika Bearybashe Bengali Bhili Bhojpuri Bundeli Chakhesang Chakma Chhattisgarhi English Garhwali Garo Gondi Gujarati Halbi Haryanvi Hindi IduMishmi Kannada Kashmiri Karbi Khariboli Khortha Kokborok Konkani… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-transcription-part.audioautomatic-speech-recognition1M<n<10M20 likes1.6k downloads6mo agoHugging Face04ivrit-ai /audio-v2-transcriptsgated Overview This dataset provides full, machine-generated transcriptions for the entire audio-v2 dataset, containing >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license. It was released on May 18th, 2025. You can find the full list of sources in this dataset under the audio-v2 dataset's sources.txt. All files were transcribed using the process.py pipeline, performing: Frame-level VAD Machine transcription using ivrit.ai's whisper-large-v3-turbo engine with the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-transcripts.audio-classification10K<n<100K1 likes882 downloads10mo agoHugging Face05ghananlpcommunity /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes692 downloads19d agoHugging Face06ghanaopenai /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes443 downloads19d agoHugging Face07openbank-uz /youtube_transcriptions Dataset Description A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google's Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings. Use Cases Automatic Speech Recognition (ASR) for Uzbek Text-to-Speech (TTS) synthesis for Uzbek Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS) Speaker-conditioned TTS… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/youtube_transcriptions.audioautomatic-speech-recognition100K<n<1M2 likes439 downloads6mo agoHugging Face08modulate /entity-transcription-benchmark Entity Transcription Benchmark Measures whether a speech recognition system transcribes named entities correctly — as distinct from word error rate. WER weights every token equally. The tokens that matter for redaction, lookup, routing and search are proper nouns, and they are a small fraction of any transcript. A system can improve WER while getting worse at exactly the words a downstream consumer needs, and nothing in the standard evaluation will show it. 2,151 clips, 6.0… See the full description on the dataset page: https://huggingface.co/datasets/modulate/entity-transcription-benchmark.audioautomatic-speech-recognition1K<n<10K4 likes295 downloads10d agoHugging Face09yuriyvnv /synthetic_transcript_pt Portuguese Speech Dataset with Multiple Training Configurations A comprehensive Portuguese speech dataset offering three distinct training configurations for speech recognition research, each designed for different experimental scenarios and training paradigms. 🎯 Dataset Configurations Overview This dataset provides three carefully curated subsets to enable comprehensive speech recognition research: Configuration Training Data Validation Test Total Samples Use Case… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_pt.audioautomatic-speech-recognition100K<n<1M0 likes271 downloads5mo agoHugging Face10Whispering-GPT /yannick-kilcher-transcript-audio Dataset Card for "yannic-kilcher-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Yannic Kilcher. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Yannic Kilcher. Data Fields The dataset is composed by: id: Id of the youtube video. channel:… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/yannick-kilcher-transcript-audio.audioautomatic-speech-recognitionn<1K0 likes268 downloads4y agoHugging Face11POTOMITAN /potomitan-gcf-transcriptiongated Kreyol Guadeloupe Transcription Dataset Ce jeu de données contient des segments audio courts (~5 secondes) en créole guadeloupéen (gcf), extraits d’émissions de radio et de télévision. Il vise à entraîner des modèles de reconnaissance automatique de la parole (ASR) pour une langue vivante mais peu disposant de peu de ressources écrites. Dataset Description Le créole guadeloupéen (Karukéya) est une langue créole à base lexicale française, parlée principalement en… See the full description on the dataset page: https://huggingface.co/datasets/POTOMITAN/potomitan-gcf-transcription.audioautomatic-speech-recognition10K<n<100K4 likes226 downloads10mo agoHugging Face12WhissleAI /indicvoices_hi_tagged_transcripts Dataset Card for indicvoices_hi_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_hi_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes220 downloads2y agoHugging Face13united-nations /transcription-corpus UN Transcription Corpus Two splits of UN meeting audio paired with official verbatim records. Splits sessions — Whole meeting sessions (SC + GA plenary) One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org. Column Description symbol UN document symbol, e.g. S/PV.9826 webtv_url URL on UN Web TV duration_ms Session duration in milliseconds num_speakers Number of speaker turns in the verbatim record audio_floor Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.audioautomatic-speech-recognitionn<1K0 likes191 downloads7mo agoHugging Face14yuriyvnv /synthetic_transcript_nl Dutch Synthetic Speech Transcripts This dataset contains 34,898 synthetic Dutch speech samples generated using GPT-4o-mini for transcript creation and OpenAI's TTS-1 model for speech synthesis. It was designed to augment Automatic Speech Recognition (ASR) training for low-resource scenarios, matching the linguistic distribution of Common Voice 17.0 Dutch. Dataset Description Purpose This dataset addresses the challenge of limited labeled speech data for Dutch… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_nl.audioautomatic-speech-recognition10K<n<100K0 likes178 downloads10mo agoHugging Face15tech4humans /Audio-Transcription-Models-Comparison-PT-BR Audio Transcription Models Comparison A dataset dedicated to comparing the performance of modern Speech-to-Text (STT) models, focusing exclusively on Brazilian Portuguese. About the Dataset This dataset was created to store and compare transcription results from different Artificial Intelligence models in challenging scenarios. Unlike generic benchmarks, this project focuses on the reality of usage in Brazil, covering: Regionalism: Local vocabulary, accents, and… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/Audio-Transcription-Models-Comparison-PT-BR.audioautomatic-speech-recognitionn<1K3 likes159 downloads7mo agoHugging Face16WhissleAI /indicvoices_pa_tagged_transcripts Dataset Card for indicvoices_pa_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_pa_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes145 downloads2y agoHugging Face17twangodev /radiotalk-us-transcripts-grok-4.20-50k radiotalk-us-transcripts-grok-4.20-50k 49,984 synthetic US air-traffic-control transcripts, generated with xAI's grok-4.20-0309-non-reasoning against the v2 radiotalk scenario pipeline. Built for fine-tuning ATC ASR models (NVIDIA Parakeet, Whisper, etc.) and for seeding TTS audio generation. Third release in the radiotalk transcripts series, and the first from a non-Qwen generator: v1: twangodev/radiotalk-us-transcripts-qwen3-100k v2:… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.20-50k.textautomatic-speech-recognition10K<n<100K0 likes127 downloads1mo agoHugging Face18AIxBlock /doctor-patient-convers-transcriptions-PII-redactedThis dataset contains real-world transcriptions of doctor–patient conversations in English (USA accent), focused on two medical specialties: ENT (Ear, Nose, Throat) and Dermatology and Orthopaedic. All conversations were originally recorded in clinical settings and transcribed by human experts. To comply with privacy regulations, only the transcription files are released, with all personally identifiable information (PII) fully redacted. Due to regulations, we are only able to publish… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/doctor-patient-convers-transcriptions-PII-redacted.automatic-speech-recognition0 likes123 downloads1y agoHugging Face19twangodev /radiotalk-us-transcripts-grok-4.3-25k radiotalk-us-transcripts-grok-4.3-25k 24,995 synthetic US air-traffic-control transcripts, generated with xAI's grok-4.3 (reasoning) against the same v2 radiotalk scenario pipeline as the earlier releases. Fourth release in the series: v1: twangodev/radiotalk-us-transcripts-qwen3-100k v2: twangodev/radiotalk-us-transcripts-qwen3-25k v3: twangodev/radiotalk-us-transcripts-grok-4.20-50k v4: this dataset Same scenario machinery, prompt p2, taxonomy t1, and realism validator as… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.3-25k.textautomatic-speech-recognition10K<n<100K0 likes117 downloads1mo agoHugging Face20twangodev /radiotalk-us-transcripts-qwen3-100k radiotalk-us-transcripts-qwen3-100k 100,000 synthetic US air-traffic-control transcripts, generated with Qwen/Qwen3-32B-NVFP4 (v1 radiotalk pipeline). First release in the radiotalk transcripts series; the v2 release with higher per-transcript realism lives at twangodev/radiotalk-us-transcripts-qwen3-25k. Renamed from radiotalk-us-transcripts-100k on 2026-08-08 to record the generator model in the dataset name; the old id redirects here. textautomatic-speech-recognition10K<n<100K0 likes104 downloads1mo agoHugging Face21twangodev /radiotalk-us-transcripts-qwen3-25k radiotalk-us-transcripts-qwen3-25k 22,065 synthetic US air-traffic-control transcripts, generated with Qwen/Qwen3-32B-NVFP4 against the v2 radiotalk pipeline. Built for fine-tuning ATC ASR models (NVIDIA Parakeet, Whisper, etc.) and for seeding TTS audio generation. This is the second release in the radiotalk transcripts series. The v1 release lives at twangodev/radiotalk-us-transcripts-qwen3-100k. What's new vs v1 v2 rebuilds the pipeline end-to-end. Lower row… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-qwen3-25k.textautomatic-speech-recognition10K<n<100K0 likes89 downloads1mo agoHugging Face22thepowerfuldeez /massive-yt-edu-transcriptions Massive YouTube Educational Transcriptions Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5. Stats Videos: 59,355 Characters: 1,539,022,925 (~384M tokens) Audio hours: 35,890 Model: faster-whisper (CTranslate2) with distil-large-v3.5 Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime Fields Field Description video_id YouTube video ID title Video title text Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.tabularautomatic-speech-recognition10K<n<100K3 likes82 downloads4mo agoHugging Face23WhissleAI /indicvoices_bn_tagged_transcripts Dataset Card for indicvoices_bn_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_bn_tagged_transcripts.audioautomatic-speech-recognitionn<1K0 likes79 downloads2y agoHugging Face24Whispering-GPT /yannic-kilcher-transcript Dataset Card for "yannic-kilcher-transcript" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Yannic Kilcher. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Yannic Kilcher. Data Fields The dataset is composed by: id: Id of the youtube video. channel: Name… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/yannic-kilcher-transcript.textautomatic-speech-recognitionn<1K1 likes64 downloads4y agoHugging Face25serge-wilson /wolof_speech_transcription Wolof Speech Transcription Description Dataset de reconnaissance automatique de la parole (ASR) en wolof, une langue d'Afrique de l'Ouest parlée par plus de 10 millions de locuteurs, principalement au Sénégal. Ce dataset est un miroir HuggingFace du corpus wolof du projet ALFFA hébergé à l'origine sur GitHub par le laboratoire GETALP (Grenoble). Source originale Ce dataset provient du projet ALFFA : Repository : getalp/ALFFA_PUBLIC Laboratoire : GETALP… See the full description on the dataset page: https://huggingface.co/datasets/serge-wilson/wolof_speech_transcription.audioautomatic-speech-recognition10K<n<100K3 likes61 downloads6mo agoHugging Face26Hk0019 /lex-fridman-podcast-transcript-audio Dataset Card for "lexFridmanPodcast-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast. Data Fields The dataset is composed by: id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Hk0019/lex-fridman-podcast-transcript-audio.audioautomatic-speech-recognitionn<1K0 likes56 downloads9mo agoHugging Face27Whispering-GPT /whisper-transcripts-ml-street-talk Dataset Card for "whisper-transcripts-mlst" More Information needed textautomatic-speech-recognitionn<1K1 likes42 downloads4y agoHugging Face28AIxBlock /Eng-Filipino-Accented-audio-with-human-transcription-call-center-topicThis dataset contains 103+ hours of spontaneous English conversations spoken in a Filipino accent, recorded in a studio environment to ensure crystal-clear audio quality. The conversations are designed as role-play scenarios between agents and customers across a variety of call center domains. 🗣️ Speech Style: Natural, unscripted role-playing between native Filipino-accented English speakers, simulating real-world customer interactions. 🎧 Audio Format: High-quality stereo WAV files, recorded… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Eng-Filipino-Accented-audio-with-human-transcription-call-center-topic.audioautomatic-speech-recognitionn<1K5 likes41 downloads1y agoHugging Face29hiraki /seamless-interact-canary-transcripts Seamless Interact - Canary Transcripts Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries. Dataset Description This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying the… See the full description on the dataset page: https://huggingface.co/datasets/hiraki/seamless-interact-canary-transcripts.tabularautomatic-speech-recognition1M<n<10M0 likes38 downloads6mo agoHugging Face30ketav /hindi-youtube-asr-transcripts Hindi YouTube ASR Transcripts Auto-generated YouTube transcripts (VTT) from 21 Hindi channels for training ASR and TTS models. Quick Start # Download and extract wget https://huggingface.co/datasets/ketav/hindi-youtube-asr-transcripts/resolve/main/youtube_asr_data.tar.gz tar -xzf youtube_asr_data.tar.gz Stats Metric Value Channels 21 Total videos 109,981 Total hours 22,186.3 Hindi subtitles 98,309 Usable hours 19,032.5 Period 2025-2026… See the full description on the dataset page: https://huggingface.co/datasets/ketav/hindi-youtube-asr-transcripts.automatic-speech-recognition10K<n<100K0 likes35 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.