datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
linustechtips-transcript-audio
Dataset Card for "linustechtips"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips.
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.lex-fridman-podcast-transcript-audio
Dataset Card for "lexFridmanPodcast-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.
Data Fields
The dataset is composed by:
id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.Malaysian-STT-Whisper
Malaysian STT Whisper format
Heavy postprocessing and post-translation to improve pseudolabeled Whisper Large V3. Also include word level timestamp.
Postprocessing
Check repetitive trigrams.
Verify Voice Activity using Silero-VAD.
Verify scores using Force Alignment.
Post-translation
We use mesolitica/nanot5-base-malaysian-translation-v2.1.
Dataset involved
Malaysian context v2
Singaporean context
Indonesian context
Mandarin audio
Tamil audio… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper.pseudolabel-malaysian-youtube-whisper-large-v3
Pseudolabel Malaysian Youtube videos using Whisper Large V3
Original dataset at https://huggingface.co/datasets/malaysia-ai/crawl-youtube, distributed pseudolabelled using 4x A100s
script at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text-semisupervised/pseudolabel-whisper
Each audio is 30 seconds.
Each audio saved in 16k sample rate.
german-asr-mixed-whisper
Dataset Card
Dataset Sources and Licensing
This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use.
Dataset Name
Original Source / Author
Link
TUDA-De German Speech Corpus
LT Group at UHH / TU Darmstadt
https://huggingface.co/datasets/uhhlt/Tuda-De
Mozilla Common Voice
Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/fosple/german-asr-mixed-whisper.common_voice_13_0-timestamped
Distil Whisper: Common Voice 13 With Timestamps
This is a variant of the Common Voice 13 dataset, augmented to return the pseudo-labelled Whisper
Transcriptions alongside the original dataset elements. The pseudo-labelled transcriptions were generated by
labelling the input audio data with the Whisper large-v2
model with greedy sampling and timestamp prediction. For information on how the original dataset was curated, refer to the original
dataset card.
Standalone Usage… See the full description on the dataset page: https://huggingface.co/datasets/distil-whisper/common_voice_13_0-timestamped.whisper-dataset-ytb-uk
Dataset Card for Dataset Name
This dataset is collected from youtube.
knesset-plenums-whisper-training
Dataset Card for ivrit.ai - Knesset Plenums Whisper Training
This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset.
This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less.
Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription.
The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.uyghur-whisper-finetune
UyZh-FolkSpeech Whisper 微调数据集
维吾尔语-汉语平行语音数据集,适用于 Whisper 模型微调。
数据集来源
本数据集源自 UyZh-FolkSpeech,经过以下处理:
音频格式转换: M4A → WAV (16kHz, 单声道)
文本清洗: 移除不可见控制字符 (U+200E, U+200F 等)
数据集划分: train/val/test (80%/10%/10%)
目录重组: 按划分分文件夹存储
数据统计
划分
记录数
音频数
时长
train
6,348
6,348
407.60 分钟
validation
793
793
48.89 分钟
test
795
795
51.80 分钟
总计
7,936
7,936
508.29 分钟
内容分布
短句 (sentence): 3,812 条
词汇短语 (word): 4,124 条
说话人分布
每个文本条目由 4 位说话人… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-whisper-finetune.Whisper-Hallucination
Whisper Hallucination and Repetition Probes
This is a BENCHMARK. Every evaluation config is test — do not fine-tune on it.
(The one exception is lexicon_synth, which is synthetic training material and ships its
own train/test split. It is not one of the eight benchmark arms — see below.)
Training on these clips invalidates every number you would then report. Build training
data separately from the same source corpora, excluding the items listed in
benchmark/exclusions.json in… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Whisper-Hallucination.whispers_win
The best whispers for Windows
Support me: Boosty or Donationalerts
yannick-kilcher-transcript-audio
Dataset Card for "yannic-kilcher-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Yannic Kilcher. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Yannic Kilcher.
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel:… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/yannick-kilcher-transcript-audio.german-asr-mixed-whisper
Dataset Card
Dataset Sources and Licensing
This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use.
Dataset Name
Original Source / Author
Link
TUDA-De German Speech Corpus
LT Group at UHH / TU Darmstadt
https://huggingface.co/datasets/uhhlt/Tuda-De
Mozilla Common Voice
Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/german-asr-mixed-whisper.tedliumThe TED-LIUM corpus is English-language TED talks, with transcriptions, sampled at 16kHz. It contains about 118 hours of speech.whisper-dari-16k
Whisper Dari 16 kHz
This dataset pairs 16 kHz WAV audio with Dari transcriptions and is arranged
for the Hugging Face audiofolder loader. The columns are audio,
transcription, and id.
from datasets import load_dataset
dataset = load_dataset("audiofolder", data_dir="whisper-dari/huggingface_dataset")
print(dataset)
The split sizes are 882 train, 111 validation, and 110 test samples.
Whispered
Time-aligned multilingual ASR enrichment over Common Voice 17
A time-aligned, quality-scored enrichment layer over Common Voice 17
for 11 languages across 7 writing systems. Each row is one Common Voice clip with:
the human transcript (the ground-truth target, used as-is),
a whisper-large-v3 machine transcript (enrichment / agreement signal — not a replacement),
word- and segment-level timestamps from MMS forced alignment of the human transcript,
language-ID, WER/CER agreement… See the full description on the dataset page: https://huggingface.co/datasets/burakaydinofficial/Whispered.raw-speech-whispervq-v1
Dataset Overview
This dataset contains over 2,4M English ASR samples, using:
The a training set of parler-tts/mls_eng_10k
Tokenized using WhisperVQ.
Usage
from datasets import load_dataset, Audio
# Load Instruction Speech dataset
dataset = load_dataset("homebrewltd/raw-speech-whispervq-v1",split='train')
Dataset Fields
Field
Type
Description
tokens
sequence
Tokenized using Encodec
text
sequence
Converted audio tokens
Bias, Risks… See the full description on the dataset page: https://huggingface.co/datasets/Menlo/raw-speech-whispervq-v1.common_voice_13_0
Distil Whisper: Common Voice 13
This is a variant of the Common Voice 13 dataset, augmented to return the pseudo-labelled Whisper
Transcriptions alongside the original dataset elements. The pseudo-labelled transcriptions were generated by
labelling the input audio data with the Whisper large-v2
model with greedy sampling. For information on how the original dataset was curated, refer to the original
dataset card.
Standalone Usage
First, install the latest version… See the full description on the dataset page: https://huggingface.co/datasets/distil-whisper/common_voice_13_0.lex-fridman-podcast
Dataset Card for "lexFridmanPodcast-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.
Data Fields
The dataset is composed by:
id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast.turkish-synthetic-whisper-rounds1-4.5-355h
Turkish Synthetic Whisper Rounds 1–4.5
Archival release of the exact 199,590-record, 355.186-hour synthetic corpus used to fine-tune the final Round 4.5 Whisper Tiny and Base models.
Each row in train.jsonl references both:
training_audio: the exact clean or exactly-once postprocessed waveform used in training; and
clean_audio: its original synthetic clean waveform.
Audio is SHA-256 deduplicated and stored in deterministic tar.zst shards. Common Voice/FLEURS evaluation audio… See the full description on the dataset page: https://huggingface.co/datasets/erayyapagci/turkish-synthetic-whisper-rounds1-4.5-355h.preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2605
5bfe2d098c8486d97fac8be76d86ec9146435245
train
56:46:32
50,557
589,095
11.7
31.9
techiaith/corpws-clllc-wlga
5d00294c31c78b1d7937bb2c2bc6cc70bc18d410
clips
48:20:49
27,579… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607.jamaican-patwa-whisperA dataset of Jamaican Patwa audio recordings with transcriptions for training speech recognition models.librispeech_asrLibriSpeech is a corpus of approximately 1000 hours of read English speech with sampling rate of 16 kHz,
prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read
audiobooks from the LibriVox project, and has been carefully segmented and aligned.87crowd-recital-whisper-training
Dataset Card for ivrit.ai - Crowd Recital
Dataset Details
Dataset Description
License
The dataset is released under the ivrit.ai License, which enables broad research and commercial use.
- Full license: https://www.ivrit.ai/en/the-license/
- FAQs: https://www.ivrit.ai/en/license-faqs/
Dataset Structure
Data Fields
Each example in the dataset contains:
audio: An audio column containing:
bytes: The audio data… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-whisper-training.gigaspeech-lGigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality
labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised
and unsupervised training. Around 40,000 hours of transcribed audio is first collected from audiobooks, podcasts
and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science,
sports, etc. A new forced alignment and segmentation pipeline is proposed to create sentence segments suitable
for speech recognition training, and to filter out segments with low-quality transcription. For system training,
GigaSpeech provides five subsets of different sizes, 10h, 250h, 1000h, 2500h, and 10000h.
For our 10,000-hour XL training subset, we cap the word error rate at 4% during the filtering/validation stage,
and for all our other smaller training subsets, we cap it at 0%. The DEV and TEST evaluation sets, on the other hand,
are re-processed by professional human transcribers to ensure high transcription quality.yannic-kilcher-transcript
Dataset Card for "yannic-kilcher-transcript"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Yannic Kilcher. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Yannic Kilcher.
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/yannic-kilcher-transcript.pseudolabel-malaya-speech-stt-train-whisper-large-v3ami-sdmThe AMI Meeting Corpus consists of 100 hours of meeting recordings. The recordings use a range of signals
synchronized to a common timeline. These include close-talking and far-field microphones, individual and
room-view video cameras, and output from a slide projector and an electronic whiteboard. During the meetings,
the participants also have unsynchronized pens available to them that record what is written. The meetings
were recorded in English using three different rooms with different acoustic properties, and include mostly
non-native speakers. \narabic-whisper-multidialect
Arabic Whisper Multi-Dialect ASR Dataset
A comprehensive multi-dialect Arabic speech recognition dataset prepared for Whisper model fine-tuning.
Dataset Description
This dataset combines high-quality Arabic speech data from multiple dialects, specifically curated for fine-tuning OpenAI's Whisper models on Arabic speech recognition tasks.
Dialects Included
Modern Standard Arabic (MSA) - Formal Arabic used in media and formal contexts
Egyptian Arabic (EGY) - The… See the full description on the dataset page: https://huggingface.co/datasets/MadLook/arabic-whisper-multidialect.spgispeechThe SPGISpeech corpus is derived from company earnings calls manually transcribed by S&P Global, Inc. according to a pro- fessional style guide detailing conventions for capitalization, punctuation, denormalization of non-standard words and tran- scription of disfluencies in spontaneous speech. The basic unit of SPGISpeech is a pair consisting of a 5 to 15 second long 16 bit, 16kHz mono wav audio file and its transcription..
