datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
linustechtips-transcript-audio
Dataset Card for "linustechtips"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips.
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.lex-fridman-podcast-transcript-audio
Dataset Card for "lexFridmanPodcast-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.
Data Fields
The dataset is composed by:
id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.Vaani-transcription-partThis dataset is part of the Vaani dataset and consists of only transcribed speech data. It has a total duration of 2041.54 hours, covering 59 languages.
This table represents the audio and transcription duration data for various languages.
Language
Angami
Angika
Ao
Assamese
Awadhi
Bajjika
Bearybashe
Bengali
Bhili
Bhojpuri
Bundeli
Chakhesang
Chakma
Chhattisgarhi
English
Garhwali
Garo
Gondi
Gujarati
Halbi
Haryanvi
Hindi
IduMishmi
Kannada
Kashmiri
Karbi
Khariboli
Khortha
Kokborok
Konkani… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-transcription-part.kumawood-speech-transcriptions
Kumawood Speech Transcriptions
Speech segments from Ghanaian films, each paired with the film's human-authored
English subtitle and a machine Twi transcript.
Total number of hours 249.8 hours
Fields
field
meaning
audio
16 kHz mono FLAC segment
text
English subtitle displayed during the segment (human-authored, recovered by OCR)
twi_text
Twi transcript from Google STT (ak) — machine output
twi_words_per_sec
transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/kumawood-speech-transcriptions.kumawood-speech-transcriptions
Kumawood Speech Transcriptions
Speech segments from Ghanaian films, each paired with the film's human-authored
English subtitle and a machine Twi transcript.
Total number of hours 249.8 hours
Fields
field
meaning
audio
16 kHz mono FLAC segment
text
English subtitle displayed during the segment (human-authored, recovered by OCR)
twi_text
Twi transcript from Google STT (ak) — machine output
twi_words_per_sec
transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kumawood-speech-transcriptions.youtube_transcriptions
Dataset Description
A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google's Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings.
Use Cases
Automatic Speech Recognition (ASR) for Uzbek
Text-to-Speech (TTS) synthesis for Uzbek
Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS)
Speaker-conditioned TTS… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/youtube_transcriptions.entity-transcription-benchmark
Entity Transcription Benchmark
Measures whether a speech recognition system transcribes named entities
correctly — as distinct from word error rate.
WER weights every token equally. The tokens that matter for redaction, lookup,
routing and search are proper nouns, and they are a small fraction of any
transcript. A system can improve WER while getting worse at exactly the words a
downstream consumer needs, and nothing in the standard evaluation will show it.
2,151 clips, 6.0… See the full description on the dataset page: https://huggingface.co/datasets/modulate/entity-transcription-benchmark.synthetic_transcript_pt
Portuguese Speech Dataset with Multiple Training Configurations
A comprehensive Portuguese speech dataset offering three distinct training configurations for speech recognition research, each designed for different experimental scenarios and training paradigms.
🎯 Dataset Configurations Overview
This dataset provides three carefully curated subsets to enable comprehensive speech recognition research:
Configuration
Training Data
Validation
Test
Total Samples
Use Case… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_pt.yannick-kilcher-transcript-audio
Dataset Card for "yannic-kilcher-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Yannic Kilcher. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Yannic Kilcher.
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel:… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/yannick-kilcher-transcript-audio.potomitan-gcf-transcription
Kreyol Guadeloupe Transcription Dataset
Ce jeu de données contient des segments audio courts (~5 secondes) en créole guadeloupéen (gcf), extraits d’émissions de radio et de télévision.
Il vise à entraîner des modèles de reconnaissance automatique de la parole (ASR) pour une langue vivante mais peu disposant de peu de ressources écrites.
Dataset Description
Le créole guadeloupéen (Karukéya) est une langue créole à base lexicale française, parlée principalement en… See the full description on the dataset page: https://huggingface.co/datasets/POTOMITAN/potomitan-gcf-transcription.transcription-corpus
UN Transcription Corpus
Two splits of UN meeting audio paired with official verbatim records.
Splits
sessions — Whole meeting sessions (SC + GA plenary)
One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org.
Column
Description
symbol
UN document symbol, e.g. S/PV.9826
webtv_url
URL on UN Web TV
duration_ms
Session duration in milliseconds
num_speakers
Number of speaker turns in the verbatim record
audio_floor
Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.synthetic_transcript_nl
Dutch Synthetic Speech Transcripts
This dataset contains 34,898 synthetic Dutch speech samples generated using GPT-4o-mini for transcript creation and OpenAI's TTS-1 model for speech synthesis. It was designed to augment Automatic Speech Recognition (ASR) training for low-resource scenarios, matching the linguistic distribution of Common Voice 17.0 Dutch.
Dataset Description
Purpose
This dataset addresses the challenge of limited labeled speech data for Dutch… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_nl.Audio-Transcription-Models-Comparison-PT-BR
Audio Transcription Models Comparison
A dataset dedicated to comparing the performance of modern Speech-to-Text (STT) models, focusing exclusively on Brazilian Portuguese.
About the Dataset
This dataset was created to store and compare transcription results from different Artificial Intelligence models in challenging scenarios. Unlike generic benchmarks, this project focuses on the reality of usage in Brazil, covering:
Regionalism: Local vocabulary, accents, and… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/Audio-Transcription-Models-Comparison-PT-BR.radiotalk-us-transcripts-grok-4.20-50k
radiotalk-us-transcripts-grok-4.20-50k
49,984 synthetic US air-traffic-control transcripts, generated with
xAI's grok-4.20-0309-non-reasoning against the v2 radiotalk scenario
pipeline. Built for fine-tuning ATC ASR models (NVIDIA Parakeet,
Whisper, etc.) and for seeding TTS audio generation.
Third release in the radiotalk transcripts series, and the first from a
non-Qwen generator:
v1: twangodev/radiotalk-us-transcripts-qwen3-100k
v2:… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.20-50k.radiotalk-us-transcripts-grok-4.3-25k
radiotalk-us-transcripts-grok-4.3-25k
24,995 synthetic US air-traffic-control transcripts, generated with xAI's
grok-4.3 (reasoning) against the same v2 radiotalk scenario pipeline as
the earlier releases. Fourth release in the series:
v1: twangodev/radiotalk-us-transcripts-qwen3-100k
v2: twangodev/radiotalk-us-transcripts-qwen3-25k
v3: twangodev/radiotalk-us-transcripts-grok-4.20-50k
v4: this dataset
Same scenario machinery, prompt p2, taxonomy t1, and realism validator
as… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.3-25k.radiotalk-us-transcripts-qwen3-25k
radiotalk-us-transcripts-qwen3-25k
22,065 synthetic US air-traffic-control transcripts, generated with
Qwen/Qwen3-32B-NVFP4 against the v2 radiotalk pipeline. Built for
fine-tuning ATC ASR models (NVIDIA Parakeet, Whisper, etc.) and for
seeding TTS audio generation.
This is the second release in the radiotalk transcripts series. The
v1 release lives at
twangodev/radiotalk-us-transcripts-qwen3-100k.
What's new vs v1
v2 rebuilds the pipeline end-to-end. Lower row… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-qwen3-25k.radiotalk-us-transcripts-qwen3-100k
radiotalk-us-transcripts-qwen3-100k
100,000 synthetic US air-traffic-control transcripts, generated with
Qwen/Qwen3-32B-NVFP4 (v1 radiotalk pipeline). First release in the
radiotalk transcripts series; the v2 release with higher per-transcript
realism lives at
twangodev/radiotalk-us-transcripts-qwen3-25k.
Renamed from radiotalk-us-transcripts-100k on 2026-08-08 to record the
generator model in the dataset name; the old id redirects here.
massive-yt-edu-transcriptions
Massive YouTube Educational Transcriptions
Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5.
Stats
Videos: 59,355
Characters: 1,539,022,925 (~384M tokens)
Audio hours: 35,890
Model: faster-whisper (CTranslate2) with distil-large-v3.5
Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime
Fields
Field
Description
video_id
YouTube video ID
title
Video title
text
Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.wolof_speech_transcription
Wolof Speech Transcription
Description
Dataset de reconnaissance automatique de la parole (ASR) en wolof, une langue d'Afrique de l'Ouest parlée par plus de 10 millions de locuteurs, principalement au Sénégal.
Ce dataset est un miroir HuggingFace du corpus wolof du projet ALFFA hébergé à l'origine sur GitHub par le laboratoire GETALP (Grenoble).
Source originale
Ce dataset provient du projet ALFFA :
Repository : getalp/ALFFA_PUBLIC
Laboratoire : GETALP… See the full description on the dataset page: https://huggingface.co/datasets/serge-wilson/wolof_speech_transcription.yannic-kilcher-transcript
Dataset Card for "yannic-kilcher-transcript"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Yannic Kilcher. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Yannic Kilcher.
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/yannic-kilcher-transcript.lex-fridman-podcast-transcript-audio
Dataset Card for "lexFridmanPodcast-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.
Data Fields
The dataset is composed by:
id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Hk0019/lex-fridman-podcast-transcript-audio.whisper-transcripts-ml-street-talk
Dataset Card for "whisper-transcripts-mlst"
More Information needed
seamless-interact-canary-transcripts
Seamless Interact - Canary Transcripts
Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries.
Dataset Description
This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying the… See the full description on the dataset page: https://huggingface.co/datasets/hiraki/seamless-interact-canary-transcripts.whisper-transcripts-the-ai-epiphany
Dataset Card for "the-ai-epiphany"
More Information needed
whisper-transcripts-linustechtips
Dataset Card for "whisper-transcripts-linustechtips"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name of the channel.
channel_id: Id of the youtube channel.
title: Title given to… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/whisper-transcripts-linustechtips.French-Medical-Transcription-Benchmark
🩺 French Medical Transcription Evaluation Dataset
Ce dataset a été créé et ouvert à la communauté dans le cadre du développement R&D de LucioleScribe, la plateforme souveraine de transcription IA 100% locale, spécifiquement conçue pour les milieux médicaux et juridiques (compatibilité RGPD, HDS, et architectures Air-Gapped).
🔗 Découvrir LucioleScribe Édition Santé | ⚙️ Voir le Pipeline Technologique Local
📊 Présentation du Dataset
L'évaluation des modèles de… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/French-Medical-Transcription-Benchmark.podcast-transcripts
Podcast Transcripts & Belief Graph
Structured belief extractions, transcripts, speaker profiles, and embeddings
mined from Bitcoin / crypto podcasts by the
be-podcast-etl pipeline.
Scale (snapshot 2026-04-21)
Asset
Count
Episodes (manifests)
1,551
Podcasts
18
Speakers
876
Persons (enriched profiles)
3,915
Belief shards
66,453
Embeddings (1536-dim)
65,007
Matrices
62,882
Top podcasts: simply-bitcoin (375), the-bitcoin-matrix (264)… See the full description on the dataset page: https://huggingface.co/datasets/BeliefEngines/podcast-transcripts.transcription-scorer
Transcription Scorer Dataset
The Transcription Scorer dataset was created to support research in reference-free evaluation of Automatic Speech Recognition (ASR) systems using human feedback. Unlike traditional evaluation metrics such as WER and its derivatives, this dataset reflects judgments of ASR outputs by human raters across multiple criteria, simulating the way a teacher grades students.
⚙️ What’s Inside
This dataset contains 1200 audio samples (from diverse sources… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/transcription-scorer.cleaned-asr-transcriptscleaned-asr-transcripts-hinglish
cleaned-asr-transcripts-hinglish
bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts.
This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.
