datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seamless-interact-canary-transcripts
Seamless Interact - Canary Transcripts
Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries.
Dataset Description
This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying the… See the full description on the dataset page: https://huggingface.co/datasets/hiraki/seamless-interact-canary-transcripts.podcast-transcripts
Podcast Transcripts Dataset
This dataset contains transcripts from Bitcoin and cryptocurrency podcasts,
processed by the belief-engines ETL pipeline.
Files
transcripts.parquet - Full episode transcripts with speaker diarization
transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings
Schema
transcripts.parquet
episode_id: Unique episode identifier
podcast_slug: Podcast name slug
episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.Podcast-Transcripts-Deduped
Podcast Transcripts (Deduped)
Phase 1 deduplicated subset of hudsongouge/Podcast-Transcripts-Raw.
Updated 2026-07-19T06-28-19Z UTC.
Dedup summary
Metric
Value
Raw input rows
102,374
Keepers
99,035
Discarded
3,339
Exact discarded
882
MinHash discarded
2457
Jaccard threshold
0.88
Two-pass Phase 1:
Exact match on canonical audio source key (YouTube ID / Omny clip / audio filename)
MinHash + LSH on normalized transcript text (Jaccard ≈… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Deduped.Podcast-Transcripts-Raw
Podcast Transcripts
Speaker-diarized transcripts from ~100k+ podcast episodes and YouTube videos (~45k hours of audio).
This dataset will be gated. Only people who are part of our team may access.
Splits
Config
Rows
Description
shows
124
Channels / podcast feeds (name, description, hosts, links)
episodes
102,374
Episode/video metadata (title, description, guests, tags, dates)
transcripts
102,374
ASR transcripts with diarized segments + YouTube… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Raw.
