datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audio-v2-transcripts
Overview
This dataset provides full, machine-generated transcriptions for the entire audio-v2 dataset, containing >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license.
It was released on May 18th, 2025.
You can find the full list of sources in this dataset under the audio-v2 dataset's sources.txt.
All files were transcribed using the process.py pipeline, performing:
Frame-level VAD
Machine transcription using ivrit.ai's whisper-large-v3-turbo engine with the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-transcripts.indicvoices_hi_tagged_transcripts
Dataset Card for indicvoices_hi_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_hi_tagged_transcripts.indicvoices_pa_tagged_transcripts
Dataset Card for indicvoices_pa_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_pa_tagged_transcripts.radiotalk-us-transcripts-grok-4.20-50k
radiotalk-us-transcripts-grok-4.20-50k
49,984 synthetic US air-traffic-control transcripts, generated with
xAI's grok-4.20-0309-non-reasoning against the v2 radiotalk scenario
pipeline. Built for fine-tuning ATC ASR models (NVIDIA Parakeet,
Whisper, etc.) and for seeding TTS audio generation.
Third release in the radiotalk transcripts series, and the first from a
non-Qwen generator:
v1: twangodev/radiotalk-us-transcripts-qwen3-100k
v2:… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.20-50k.radiotalk-us-transcripts-grok-4.3-25k
radiotalk-us-transcripts-grok-4.3-25k
24,995 synthetic US air-traffic-control transcripts, generated with xAI's
grok-4.3 (reasoning) against the same v2 radiotalk scenario pipeline as
the earlier releases. Fourth release in the series:
v1: twangodev/radiotalk-us-transcripts-qwen3-100k
v2: twangodev/radiotalk-us-transcripts-qwen3-25k
v3: twangodev/radiotalk-us-transcripts-grok-4.20-50k
v4: this dataset
Same scenario machinery, prompt p2, taxonomy t1, and realism validator
as… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.3-25k.radiotalk-us-transcripts-qwen3-100k
radiotalk-us-transcripts-qwen3-100k
100,000 synthetic US air-traffic-control transcripts, generated with
Qwen/Qwen3-32B-NVFP4 (v1 radiotalk pipeline). First release in the
radiotalk transcripts series; the v2 release with higher per-transcript
realism lives at
twangodev/radiotalk-us-transcripts-qwen3-25k.
Renamed from radiotalk-us-transcripts-100k on 2026-08-08 to record the
generator model in the dataset name; the old id redirects here.
radiotalk-us-transcripts-qwen3-25k
radiotalk-us-transcripts-qwen3-25k
22,065 synthetic US air-traffic-control transcripts, generated with
Qwen/Qwen3-32B-NVFP4 against the v2 radiotalk pipeline. Built for
fine-tuning ATC ASR models (NVIDIA Parakeet, Whisper, etc.) and for
seeding TTS audio generation.
This is the second release in the radiotalk transcripts series. The
v1 release lives at
twangodev/radiotalk-us-transcripts-qwen3-100k.
What's new vs v1
v2 rebuilds the pipeline end-to-end. Lower row… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-qwen3-25k.indicvoices_bn_tagged_transcripts
Dataset Card for indicvoices_bn_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_bn_tagged_transcripts.whisper-transcripts-ml-street-talk
Dataset Card for "whisper-transcripts-mlst"
More Information needed
seamless-interact-canary-transcripts
Seamless Interact - Canary Transcripts
Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries.
Dataset Description
This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying the… See the full description on the dataset page: https://huggingface.co/datasets/hiraki/seamless-interact-canary-transcripts.hindi-youtube-asr-transcripts
Hindi YouTube ASR Transcripts
Auto-generated YouTube transcripts (VTT) from 21 Hindi channels for training ASR and TTS models.
Quick Start
# Download and extract
wget https://huggingface.co/datasets/ketav/hindi-youtube-asr-transcripts/resolve/main/youtube_asr_data.tar.gz
tar -xzf youtube_asr_data.tar.gz
Stats
Metric
Value
Channels
21
Total videos
109,981
Total hours
22,186.3
Hindi subtitles
98,309
Usable hours
19,032.5
Period
2025-2026… See the full description on the dataset page: https://huggingface.co/datasets/ketav/hindi-youtube-asr-transcripts.whisper-transcripts-linustechtips
Dataset Card for "whisper-transcripts-linustechtips"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name of the channel.
channel_id: Id of the youtube channel.
title: Title given to… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/whisper-transcripts-linustechtips.whisper-transcripts-the-ai-epiphany
Dataset Card for "the-ai-epiphany"
More Information needed
podcast-transcripts
Podcast Transcripts & Belief Graph
Structured belief extractions, transcripts, speaker profiles, and embeddings
mined from Bitcoin / crypto podcasts by the
be-podcast-etl pipeline.
Scale (snapshot 2026-04-21)
Asset
Count
Episodes (manifests)
1,551
Podcasts
18
Speakers
876
Persons (enriched profiles)
3,915
Belief shards
66,453
Embeddings (1536-dim)
65,007
Matrices
62,882
Top podcasts: simply-bitcoin (375), the-bitcoin-matrix (264)… See the full description on the dataset page: https://huggingface.co/datasets/BeliefEngines/podcast-transcripts.indicvoices_mr_tagged_transcripts
Dataset Card for indicvoices_mr_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_mr_tagged_transcripts.cleaned-asr-transcriptsgigaspeech2-vi-missing-transcripts
GigaSpeech2 Vietnamese WAVs Missing Transcripts
This dataset contains the Vietnamese GigaSpeech2 training WAV files whose IDs
are absent from train_refined.tsv.
Contents
201,295 WAV files without a matching transcript
41 uncompressed TAR shards in shards/
36 GB of audio (approximately)
processing_manifest.jsonl: per-source-archive counts
summary.json: aggregate counts
SHA256SUMS: checksums for all TAR shards
Each shard preserves the original relative layout:… See the full description on the dataset page: https://huggingface.co/datasets/quangdung/gigaspeech2-vi-missing-transcripts.cleaned-asr-transcripts-hinglish
cleaned-asr-transcripts-hinglish
bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts.
This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.cleaned-asr-transcripts-hinglish
cleaned-asr-transcripts-hinglish
bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts.
This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.github_fetch_huggingface_pdf-tools_terminal_2096-docaudit-7c91-speech-transcripts
Speech Transcripts
Dataset Summary
Time-aligned transcripts of English speech audio.
Dataset Structure
Data fields: audio_path, text, start_time, end_time.
Licensing Information
This dataset is released under the CC BY-SA 4.0 license (cc-by-sa-4.0).
podcast-transcripts
Podcast Transcripts Dataset
This dataset contains transcripts from Bitcoin and cryptocurrency podcasts,
processed by the belief-engines ETL pipeline.
Files
transcripts.parquet - Full episode transcripts with speaker diarization
transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings
Schema
transcripts.parquet
episode_id: Unique episode identifier
podcast_slug: Podcast name slug
episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.goatis-transcripts
Goatis / Sv3rige Video Transcripts
Full transcripts of 1,383 videos (~487 hours, ~3.9M words) from the
YouTube channels of Goatis (Sv3rige) — the sv3rige channel (2011–2025) and the
current Goatis channel (2019–2026). This is the dataset behind
goatis.net, a searchable archive in the style of
aajonus.net.
What makes it more than raw ASR
Every video was processed with speaker identification, not just
transcription. He mostly reacts to other people's videos, so a… See the full description on the dataset page: https://huggingface.co/datasets/exoarbuus/goatis-transcripts.Podcast-Transcripts-Deduped
Podcast Transcripts (Deduped)
Phase 1 deduplicated subset of hudsongouge/Podcast-Transcripts-Raw.
Updated 2026-07-19T06-28-19Z UTC.
Dedup summary
Metric
Value
Raw input rows
102,374
Keepers
99,035
Discarded
3,339
Exact discarded
882
MinHash discarded
2457
Jaccard threshold
0.88
Two-pass Phase 1:
Exact match on canonical audio source key (YouTube ID / Omny clip / audio filename)
MinHash + LSH on normalized transcript text (Jaccard ≈… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Deduped.Podcast-Transcripts-Raw
Podcast Transcripts
Speaker-diarized transcripts from ~100k+ podcast episodes and YouTube videos (~45k hours of audio).
This dataset will be gated. Only people who are part of our team may access.
Splits
Config
Rows
Description
shows
124
Channels / podcast feeds (name, description, hosts, links)
episodes
102,374
Episode/video metadata (title, description, guests, tags, dates)
transcripts
102,374
ASR transcripts with diarized segments + YouTube… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Raw.
