datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
radiotalk-us-transcripts-grok-4.20-50k
radiotalk-us-transcripts-grok-4.20-50k
49,984 synthetic US air-traffic-control transcripts, generated with
xAI's grok-4.20-0309-non-reasoning against the v2 radiotalk scenario
pipeline. Built for fine-tuning ATC ASR models (NVIDIA Parakeet,
Whisper, etc.) and for seeding TTS audio generation.
Third release in the radiotalk transcripts series, and the first from a
non-Qwen generator:
v1: twangodev/radiotalk-us-transcripts-qwen3-100k
v2:… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.20-50k.radiotalk-us-transcripts-grok-4.3-25k
radiotalk-us-transcripts-grok-4.3-25k
24,995 synthetic US air-traffic-control transcripts, generated with xAI's
grok-4.3 (reasoning) against the same v2 radiotalk scenario pipeline as
the earlier releases. Fourth release in the series:
v1: twangodev/radiotalk-us-transcripts-qwen3-100k
v2: twangodev/radiotalk-us-transcripts-qwen3-25k
v3: twangodev/radiotalk-us-transcripts-grok-4.20-50k
v4: this dataset
Same scenario machinery, prompt p2, taxonomy t1, and realism validator
as… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.3-25k.radiotalk-us-transcripts-qwen3-25k
radiotalk-us-transcripts-qwen3-25k
22,065 synthetic US air-traffic-control transcripts, generated with
Qwen/Qwen3-32B-NVFP4 against the v2 radiotalk pipeline. Built for
fine-tuning ATC ASR models (NVIDIA Parakeet, Whisper, etc.) and for
seeding TTS audio generation.
This is the second release in the radiotalk transcripts series. The
v1 release lives at
twangodev/radiotalk-us-transcripts-qwen3-100k.
What's new vs v1
v2 rebuilds the pipeline end-to-end. Lower row… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-qwen3-25k.radiotalk-us-transcripts-qwen3-100k
radiotalk-us-transcripts-qwen3-100k
100,000 synthetic US air-traffic-control transcripts, generated with
Qwen/Qwen3-32B-NVFP4 (v1 radiotalk pipeline). First release in the
radiotalk transcripts series; the v2 release with higher per-transcript
realism lives at
twangodev/radiotalk-us-transcripts-qwen3-25k.
Renamed from radiotalk-us-transcripts-100k on 2026-08-08 to record the
generator model in the dataset name; the old id redirects here.
whisper-transcripts-ml-street-talk
Dataset Card for "whisper-transcripts-mlst"
More Information needed
whisper-transcripts-the-ai-epiphany
Dataset Card for "the-ai-epiphany"
More Information needed
whisper-transcripts-linustechtips
Dataset Card for "whisper-transcripts-linustechtips"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name of the channel.
channel_id: Id of the youtube channel.
title: Title given to… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/whisper-transcripts-linustechtips.seamless-interact-canary-transcripts
Seamless Interact - Canary Transcripts
Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries.
Dataset Description
This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying the… See the full description on the dataset page: https://huggingface.co/datasets/hiraki/seamless-interact-canary-transcripts.podcast-transcripts
Podcast Transcripts & Belief Graph
Structured belief extractions, transcripts, speaker profiles, and embeddings
mined from Bitcoin / crypto podcasts by the
be-podcast-etl pipeline.
Scale (snapshot 2026-04-21)
Asset
Count
Episodes (manifests)
1,551
Podcasts
18
Speakers
876
Persons (enriched profiles)
3,915
Belief shards
66,453
Embeddings (1536-dim)
65,007
Matrices
62,882
Top podcasts: simply-bitcoin (375), the-bitcoin-matrix (264)… See the full description on the dataset page: https://huggingface.co/datasets/BeliefEngines/podcast-transcripts.podcast-transcripts
Podcast Transcripts Dataset
This dataset contains transcripts from Bitcoin and cryptocurrency podcasts,
processed by the belief-engines ETL pipeline.
Files
transcripts.parquet - Full episode transcripts with speaker diarization
transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings
Schema
transcripts.parquet
episode_id: Unique episode identifier
podcast_slug: Podcast name slug
episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.Podcast-Transcripts-Deduped
Podcast Transcripts (Deduped)
Phase 1 deduplicated subset of hudsongouge/Podcast-Transcripts-Raw.
Updated 2026-07-19T06-28-19Z UTC.
Dedup summary
Metric
Value
Raw input rows
102,374
Keepers
99,035
Discarded
3,339
Exact discarded
882
MinHash discarded
2457
Jaccard threshold
0.88
Two-pass Phase 1:
Exact match on canonical audio source key (YouTube ID / Omny clip / audio filename)
MinHash + LSH on normalized transcript text (Jaccard ≈… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Deduped.Podcast-Transcripts-Raw
Podcast Transcripts
Speaker-diarized transcripts from ~100k+ podcast episodes and YouTube videos (~45k hours of audio).
This dataset will be gated. Only people who are part of our team may access.
Splits
Config
Rows
Description
shows
124
Channels / podcast feeds (name, description, hosts, links)
episodes
102,374
Episode/video metadata (title, description, guests, tags, dates)
transcripts
102,374
ASR transcripts with diarized segments + YouTube… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Raw.
