CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ivrit-ai /audio-v2-transcriptsgated Overview This dataset provides full, machine-generated transcriptions for the entire audio-v2 dataset, containing >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license. It was released on May 18th, 2025. You can find the full list of sources in this dataset under the audio-v2 dataset's sources.txt. All files were transcribed using the process.py pipeline, performing: Frame-level VAD Machine transcription using ivrit.ai's whisper-large-v3-turbo engine with the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-transcripts.audio-classification10K<n<100K1 likes882 downloads10mo agoHugging Face02WhissleAI /indicvoices_hi_tagged_transcripts Dataset Card for indicvoices_hi_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_hi_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes220 downloads2y agoHugging Face03WhissleAI /indicvoices_pa_tagged_transcripts Dataset Card for indicvoices_pa_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_pa_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes145 downloads2y agoHugging Face04twangodev /radiotalk-us-transcripts-grok-4.20-50k radiotalk-us-transcripts-grok-4.20-50k 49,984 synthetic US air-traffic-control transcripts, generated with xAI's grok-4.20-0309-non-reasoning against the v2 radiotalk scenario pipeline. Built for fine-tuning ATC ASR models (NVIDIA Parakeet, Whisper, etc.) and for seeding TTS audio generation. Third release in the radiotalk transcripts series, and the first from a non-Qwen generator: v1: twangodev/radiotalk-us-transcripts-qwen3-100k v2:… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.20-50k.textautomatic-speech-recognition10K<n<100K0 likes127 downloads1mo agoHugging Face05twangodev /radiotalk-us-transcripts-grok-4.3-25k radiotalk-us-transcripts-grok-4.3-25k 24,995 synthetic US air-traffic-control transcripts, generated with xAI's grok-4.3 (reasoning) against the same v2 radiotalk scenario pipeline as the earlier releases. Fourth release in the series: v1: twangodev/radiotalk-us-transcripts-qwen3-100k v2: twangodev/radiotalk-us-transcripts-qwen3-25k v3: twangodev/radiotalk-us-transcripts-grok-4.20-50k v4: this dataset Same scenario machinery, prompt p2, taxonomy t1, and realism validator as… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.3-25k.textautomatic-speech-recognition10K<n<100K0 likes117 downloads1mo agoHugging Face06twangodev /radiotalk-us-transcripts-qwen3-100k radiotalk-us-transcripts-qwen3-100k 100,000 synthetic US air-traffic-control transcripts, generated with Qwen/Qwen3-32B-NVFP4 (v1 radiotalk pipeline). First release in the radiotalk transcripts series; the v2 release with higher per-transcript realism lives at twangodev/radiotalk-us-transcripts-qwen3-25k. Renamed from radiotalk-us-transcripts-100k on 2026-08-08 to record the generator model in the dataset name; the old id redirects here. textautomatic-speech-recognition10K<n<100K0 likes104 downloads1mo agoHugging Face07twangodev /radiotalk-us-transcripts-qwen3-25k radiotalk-us-transcripts-qwen3-25k 22,065 synthetic US air-traffic-control transcripts, generated with Qwen/Qwen3-32B-NVFP4 against the v2 radiotalk pipeline. Built for fine-tuning ATC ASR models (NVIDIA Parakeet, Whisper, etc.) and for seeding TTS audio generation. This is the second release in the radiotalk transcripts series. The v1 release lives at twangodev/radiotalk-us-transcripts-qwen3-100k. What's new vs v1 v2 rebuilds the pipeline end-to-end. Lower row… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-qwen3-25k.textautomatic-speech-recognition10K<n<100K0 likes89 downloads1mo agoHugging Face08WhissleAI /indicvoices_bn_tagged_transcripts Dataset Card for indicvoices_bn_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_bn_tagged_transcripts.audioautomatic-speech-recognitionn<1K0 likes79 downloads2y agoHugging Face09Whispering-GPT /whisper-transcripts-ml-street-talk Dataset Card for "whisper-transcripts-mlst" More Information needed textautomatic-speech-recognitionn<1K1 likes42 downloads4y agoHugging Face10hiraki /seamless-interact-canary-transcripts Seamless Interact - Canary Transcripts Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries. Dataset Description This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying the… See the full description on the dataset page: https://huggingface.co/datasets/hiraki/seamless-interact-canary-transcripts.tabularautomatic-speech-recognition1M<n<10M0 likes38 downloads6mo agoHugging Face11ketav /hindi-youtube-asr-transcripts Hindi YouTube ASR Transcripts Auto-generated YouTube transcripts (VTT) from 21 Hindi channels for training ASR and TTS models. Quick Start # Download and extract wget https://huggingface.co/datasets/ketav/hindi-youtube-asr-transcripts/resolve/main/youtube_asr_data.tar.gz tar -xzf youtube_asr_data.tar.gz Stats Metric Value Channels 21 Total videos 109,981 Total hours 22,186.3 Hindi subtitles 98,309 Usable hours 19,032.5 Period 2025-2026… See the full description on the dataset page: https://huggingface.co/datasets/ketav/hindi-youtube-asr-transcripts.automatic-speech-recognition10K<n<100K0 likes35 downloads6mo agoHugging Face12Whispering-GPT /whisper-transcripts-linustechtips Dataset Card for "whisper-transcripts-linustechtips" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset Data Fields The dataset is composed by: id: Id of the youtube video. channel: Name of the channel. channel_id: Id of the youtube channel. title: Title given to… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/whisper-transcripts-linustechtips.textautomatic-speech-recognition1K<n<10K2 likes33 downloads4y agoHugging Face13Whispering-GPT /whisper-transcripts-the-ai-epiphany Dataset Card for "the-ai-epiphany" More Information needed textautomatic-speech-recognitionn<1K0 likes32 downloads4y agoHugging Face14BeliefEngines /podcast-transcripts Podcast Transcripts & Belief Graph Structured belief extractions, transcripts, speaker profiles, and embeddings mined from Bitcoin / crypto podcasts by the be-podcast-etl pipeline. Scale (snapshot 2026-04-21) Asset Count Episodes (manifests) 1,551 Podcasts 18 Speakers 876 Persons (enriched profiles) 3,915 Belief shards 66,453 Embeddings (1536-dim) 65,007 Matrices 62,882 Top podcasts: simply-bitcoin (375), the-bitcoin-matrix (264)… See the full description on the dataset page: https://huggingface.co/datasets/BeliefEngines/podcast-transcripts.texttext-generationn<1K0 likes25 downloads5mo agoHugging Face15WhissleAI /indicvoices_mr_tagged_transcripts Dataset Card for indicvoices_mr_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_mr_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes22 downloads2y agoHugging Face16bingbangboom /cleaned-asr-transcriptstexttext-generation10K<n<100K1 likes21 downloads6mo agoHugging Face17quangdung /gigaspeech2-vi-missing-transcripts GigaSpeech2 Vietnamese WAVs Missing Transcripts This dataset contains the Vietnamese GigaSpeech2 training WAV files whose IDs are absent from train_refined.tsv. Contents 201,295 WAV files without a matching transcript 41 uncompressed TAR shards in shards/ 36 GB of audio (approximately) processing_manifest.jsonl: per-source-archive counts summary.json: aggregate counts SHA256SUMS: checksums for all TAR shards Each shard preserves the original relative layout:… See the full description on the dataset page: https://huggingface.co/datasets/quangdung/gigaspeech2-vi-missing-transcripts.audioautomatic-speech-recognition0 likes21 downloads2mo agoHugging Face18SaiyanSai /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes17 downloads4mo agoHugging Face19bingbangboom /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes16 downloads5mo agoHugging Face20Roy229 /github_fetch_huggingface_pdf-tools_terminal_2096-docaudit-7c91-speech-transcripts Speech Transcripts Dataset Summary Time-aligned transcripts of English speech audio. Dataset Structure Data fields: audio_path, text, start_time, end_time. Licensing Information This dataset is released under the CC BY-SA 4.0 license (cc-by-sa-4.0). automatic-speech-recognition0 likes13 downloads1mo agoHugging Face21rchiera /podcast-transcripts Podcast Transcripts Dataset This dataset contains transcripts from Bitcoin and cryptocurrency podcasts, processed by the belief-engines ETL pipeline. Files transcripts.parquet - Full episode transcripts with speaker diarization transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings Schema transcripts.parquet episode_id: Unique episode identifier podcast_slug: Podcast name slug episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.tabulartext-generation10K<n<100K0 likes12 downloads7mo agoHugging Face22exoarbuus /goatis-transcripts Goatis / Sv3rige Video Transcripts Full transcripts of 1,383 videos (~487 hours, ~3.9M words) from the YouTube channels of Goatis (Sv3rige) — the sv3rige channel (2011–2025) and the current Goatis channel (2019–2026). This is the dataset behind goatis.net, a searchable archive in the style of aajonus.net. What makes it more than raw ASR Every video was processed with speaker identification, not just transcription. He mostly reacts to other people's videos, so a… See the full description on the dataset page: https://huggingface.co/datasets/exoarbuus/goatis-transcripts.textautomatic-speech-recognition1K<n<10K0 likes10 downloads2mo agoHugging Face23hudsongouge /Podcast-Transcripts-Dedupedgated Podcast Transcripts (Deduped) Phase 1 deduplicated subset of hudsongouge/Podcast-Transcripts-Raw. Updated 2026-07-19T06-28-19Z UTC. Dedup summary Metric Value Raw input rows 102,374 Keepers 99,035 Discarded 3,339 Exact discarded 882 MinHash discarded 2457 Jaccard threshold 0.88 Two-pass Phase 1: Exact match on canonical audio source key (YouTube ID / Omny clip / audio filename) MinHash + LSH on normalized transcript text (Jaccard ≈… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Deduped.tabularautomatic-speech-recognition100K<n<1M1 likes9 downloads2mo agoHugging Face24hudsongouge /Podcast-Transcripts-Rawgated Podcast Transcripts Speaker-diarized transcripts from ~100k+ podcast episodes and YouTube videos (~45k hours of audio). This dataset will be gated. Only people who are part of our team may access. Splits Config Rows Description shows 124 Channels / podcast feeds (name, description, hosts, links) episodes 102,374 Episode/video metadata (title, description, guests, tags, dates) transcripts 102,374 ASR transcripts with diarized segments + YouTube… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Raw.tabularautomatic-speech-recognition100K<n<1M1 likes4 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.