CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hiraki /seamless-interact-canary-transcripts Seamless Interact - Canary Transcripts Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries. Dataset Description This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying the… See the full description on the dataset page: https://huggingface.co/datasets/hiraki/seamless-interact-canary-transcripts.tabularautomatic-speech-recognition1M<n<10M0 likes37 downloads6mo agoHugging Face02rchiera /podcast-transcripts Podcast Transcripts Dataset This dataset contains transcripts from Bitcoin and cryptocurrency podcasts, processed by the belief-engines ETL pipeline. Files transcripts.parquet - Full episode transcripts with speaker diarization transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings Schema transcripts.parquet episode_id: Unique episode identifier podcast_slug: Podcast name slug episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.tabulartext-generation10K<n<100K0 likes12 downloads7mo agoHugging Face03hudsongouge /Podcast-Transcripts-Dedupedgated Podcast Transcripts (Deduped) Phase 1 deduplicated subset of hudsongouge/Podcast-Transcripts-Raw. Updated 2026-07-19T06-28-19Z UTC. Dedup summary Metric Value Raw input rows 102,374 Keepers 99,035 Discarded 3,339 Exact discarded 882 MinHash discarded 2457 Jaccard threshold 0.88 Two-pass Phase 1: Exact match on canonical audio source key (YouTube ID / Omny clip / audio filename) MinHash + LSH on normalized transcript text (Jaccard ≈… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Deduped.tabularautomatic-speech-recognition100K<n<1M1 likes9 downloads2mo agoHugging Face04hudsongouge /Podcast-Transcripts-Rawgated Podcast Transcripts Speaker-diarized transcripts from ~100k+ podcast episodes and YouTube videos (~45k hours of audio). This dataset will be gated. Only people who are part of our team may access. Splits Config Rows Description shows 124 Channels / podcast feeds (name, description, hosts, links) episodes 102,374 Episode/video metadata (title, description, guests, tags, dates) transcripts 102,374 ASR transcripts with diarized segments + YouTube… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Raw.tabularautomatic-speech-recognition100K<n<1M1 likes4 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.