CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01twangodev /radiotalk-us-transcripts-grok-4.20-50k radiotalk-us-transcripts-grok-4.20-50k 49,984 synthetic US air-traffic-control transcripts, generated with xAI's grok-4.20-0309-non-reasoning against the v2 radiotalk scenario pipeline. Built for fine-tuning ATC ASR models (NVIDIA Parakeet, Whisper, etc.) and for seeding TTS audio generation. Third release in the radiotalk transcripts series, and the first from a non-Qwen generator: v1: twangodev/radiotalk-us-transcripts-qwen3-100k v2:… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.20-50k.textautomatic-speech-recognition10K<n<100K0 likes116 downloads2mo agoHugging Face02twangodev /radiotalk-us-transcripts-grok-4.3-25k radiotalk-us-transcripts-grok-4.3-25k 24,995 synthetic US air-traffic-control transcripts, generated with xAI's grok-4.3 (reasoning) against the same v2 radiotalk scenario pipeline as the earlier releases. Fourth release in the series: v1: twangodev/radiotalk-us-transcripts-qwen3-100k v2: twangodev/radiotalk-us-transcripts-qwen3-25k v3: twangodev/radiotalk-us-transcripts-grok-4.20-50k v4: this dataset Same scenario machinery, prompt p2, taxonomy t1, and realism validator as… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.3-25k.textautomatic-speech-recognition10K<n<100K0 likes103 downloads2mo agoHugging Face03twangodev /radiotalk-us-transcripts-qwen3-25k radiotalk-us-transcripts-qwen3-25k 22,065 synthetic US air-traffic-control transcripts, generated with Qwen/Qwen3-32B-NVFP4 against the v2 radiotalk pipeline. Built for fine-tuning ATC ASR models (NVIDIA Parakeet, Whisper, etc.) and for seeding TTS audio generation. This is the second release in the radiotalk transcripts series. The v1 release lives at twangodev/radiotalk-us-transcripts-qwen3-100k. What's new vs v1 v2 rebuilds the pipeline end-to-end. Lower row… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-qwen3-25k.textautomatic-speech-recognition10K<n<100K0 likes99 downloads2mo agoHugging Face04twangodev /radiotalk-us-transcripts-qwen3-100k radiotalk-us-transcripts-qwen3-100k 100,000 synthetic US air-traffic-control transcripts, generated with Qwen/Qwen3-32B-NVFP4 (v1 radiotalk pipeline). First release in the radiotalk transcripts series; the v2 release with higher per-transcript realism lives at twangodev/radiotalk-us-transcripts-qwen3-25k. Renamed from radiotalk-us-transcripts-100k on 2026-08-08 to record the generator model in the dataset name; the old id redirects here. textautomatic-speech-recognition10K<n<100K0 likes94 downloads2mo agoHugging Face05Whispering-GPT /whisper-transcripts-ml-street-talk Dataset Card for "whisper-transcripts-mlst" More Information needed textautomatic-speech-recognitionn<1K1 likes41 downloads4y agoHugging Face06Whispering-GPT /whisper-transcripts-the-ai-epiphany Dataset Card for "the-ai-epiphany" More Information needed textautomatic-speech-recognitionn<1K0 likes38 downloads4y agoHugging Face07Whispering-GPT /whisper-transcripts-linustechtips Dataset Card for "whisper-transcripts-linustechtips" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset Data Fields The dataset is composed by: id: Id of the youtube video. channel: Name of the channel. channel_id: Id of the youtube channel. title: Title given to… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/whisper-transcripts-linustechtips.textautomatic-speech-recognition1K<n<10K2 likes34 downloads4y agoHugging Face08hiraki /seamless-interact-canary-transcripts Seamless Interact - Canary Transcripts Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries. Dataset Description This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying the… See the full description on the dataset page: https://huggingface.co/datasets/hiraki/seamless-interact-canary-transcripts.tabularautomatic-speech-recognition1M<n<10M0 likes34 downloads6mo agoHugging Face09BeliefEngines /podcast-transcripts Podcast Transcripts & Belief Graph Structured belief extractions, transcripts, speaker profiles, and embeddings mined from Bitcoin / crypto podcasts by the be-podcast-etl pipeline. Scale (snapshot 2026-04-21) Asset Count Episodes (manifests) 1,551 Podcasts 18 Speakers 876 Persons (enriched profiles) 3,915 Belief shards 66,453 Embeddings (1536-dim) 65,007 Matrices 62,882 Top podcasts: simply-bitcoin (375), the-bitcoin-matrix (264)… See the full description on the dataset page: https://huggingface.co/datasets/BeliefEngines/podcast-transcripts.texttext-generationn<1K0 likes22 downloads5mo agoHugging Face10rchiera /podcast-transcripts Podcast Transcripts Dataset This dataset contains transcripts from Bitcoin and cryptocurrency podcasts, processed by the belief-engines ETL pipeline. Files transcripts.parquet - Full episode transcripts with speaker diarization transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings Schema transcripts.parquet episode_id: Unique episode identifier podcast_slug: Podcast name slug episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.tabulartext-generation10K<n<100K0 likes12 downloads7mo agoHugging Face11hudsongouge /Podcast-Transcripts-Dedupedgated Podcast Transcripts (Deduped) Phase 1 deduplicated subset of hudsongouge/Podcast-Transcripts-Raw. Updated 2026-07-19T06-28-19Z UTC. Dedup summary Metric Value Raw input rows 102,374 Keepers 99,035 Discarded 3,339 Exact discarded 882 MinHash discarded 2457 Jaccard threshold 0.88 Two-pass Phase 1: Exact match on canonical audio source key (YouTube ID / Omny clip / audio filename) MinHash + LSH on normalized transcript text (Jaccard ≈… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Deduped.tabularautomatic-speech-recognition100K<n<1M1 likes9 downloads2mo agoHugging Face12hudsongouge /Podcast-Transcripts-Rawgated Podcast Transcripts Speaker-diarized transcripts from ~100k+ podcast episodes and YouTube videos (~45k hours of audio). This dataset will be gated. Only people who are part of our team may access. Splits Config Rows Description shows 124 Channels / podcast feeds (name, description, hosts, links) episodes 102,374 Episode/video metadata (title, description, guests, tags, dates) transcripts 102,374 ASR transcripts with diarized segments + YouTube… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Raw.tabularautomatic-speech-recognition100K<n<1M1 likes5 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.