datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
whisper_transcriptions.reazon_speech_all.wer_10.0.vectorizedwhisper_transcriptions.reazon_speech_allwhisper_transcriptions.mls.wer_10.0.vectorizedwhisper_transcriptions.mls.wer_10.0librispeech_long
Dataset Card for "librispeech_long"
More Information needed
whisper_transcriptions.reazonspeech.allwhisperspeech-librilightThis is a processed LibriLight dataset ready for training the WhisperSpeech models.
See https://github.com/collabora/WhisperSpeech for more details.
Quick start
If you want to quickly train a basic WhisperSpeech model you can start by downloading the small subset:
# magic includes to download only the small and validation data splits and the accompanying config files
huggingface-cli download --repo-type dataset --include '*-small-*' '*small.dataset' '*-speakers*' --local-dir . --… See the full description on the dataset page: https://huggingface.co/datasets/collabora/whisperspeech-librilight.librispeech_asr-noise
Dataset Card for "librispeech_asr-noise"
More Information needed
whisperjav-wheels
WhisperJAV Pre-built Wheels
Pre-built Python wheels for WhisperJAV dependencies that are difficult to compile from source.
Repository Structure
whisperjav-wheels/
├── llama-cpp-python/
│ ├── cu124/ # CUDA 12.4 wheels
│ ├── cu121/ # CUDA 12.1 wheels (legacy)
│ └── metal/ # Apple Silicon wheels
└── README.md
Available Wheels
llama-cpp-python
For local LLM translation support (whisperjav-translate --provider local).… See the full description on the dataset page: https://huggingface.co/datasets/mei986/whisperjav-wheels.whisper_transcriptions.reazonspeech.all.wer_10.0figureswhisperkit-evals
WhisperKit
WhisperKit is an on-device speech recognition framework for Apple Silicon:
https://github.com/argmaxinc/WhisperKit
For performance and accuracy benchmarks on real devices, please see:
https://huggingface.co/spaces/argmaxinc/whisperkit-benchmarks
earnings22
Dataset Card for Earnings 22
Dataset Summary
Earnings-22 provides a free-to-use benchmark of real-world, accented audio to bridge academic and industrial research.
This dataset contains 125 files totalling roughly 119 hours of English language earnings calls from global countries.
This dataset provides the full audios, transcripts, and accompanying metadata such as ticker symbol, headquarters country,
and our defined "Language Region".
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/distil-whisper/earnings22.Live-WhisperX-526K
Dataset Card for Live-WhisperX-526K
Uses
This dataset is used for the training of the LiveCC-7B-Instruct model. We only allow the use of this dataset for academic research and educational purposes. For OpenAI GPT-4o generated user prompts, we recommend users check the OpenAI Usage Policy.
Project Page: https://showlab.github.io/livecc
Paper: https://huggingface.co/papers/2504.16030
Data Sources
After we finished the pre-training of LiveCC-7B-Base… See the full description on the dataset page: https://huggingface.co/datasets/chenjoya/Live-WhisperX-526K.meanwhile
Dataset Card for "meanwhile"
This dataset consists of 64 segments from The Late Show with Stephen Colbert. This dataset was published as
part of the Whisper release by OpenAI. See page 19 of the Whisper paper
for details.
linustechtips-transcript-audio
Dataset Card for "linustechtips"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips.
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.whisper.cpp
whisper.cpp
Stable: v1.8.1 / Roadmap
High-performance inference of OpenAI's Whisper automatic speech recognition (ASR) model:
Plain C/C++ implementation without dependencies
Apple Silicon first-class citizen - optimized via ARM NEON, Accelerate framework, Metal and Core ML
AVX intrinsics support for x86 architectures
VSX intrinsics support for POWER architectures
Mixed F16 / F32 precision
Integer quantization support
Zero memory allocations at runtime
Vulkan support
Support… See the full description on the dataset page: https://huggingface.co/datasets/echodict/whisper.cpp.lex-fridman-podcast-transcript-audio
Dataset Card for "lexFridmanPodcast-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.
Data Fields
The dataset is composed by:
id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.Malaysian-STT-Whisper
Malaysian STT Whisper format
Heavy postprocessing and post-translation to improve pseudolabeled Whisper Large V3. Also include word level timestamp.
Postprocessing
Check repetitive trigrams.
Verify Voice Activity using Silero-VAD.
Verify scores using Force Alignment.
Post-translation
We use mesolitica/nanot5-base-malaysian-translation-v2.1.
Dataset involved
Malaysian context v2
Singaporean context
Indonesian context
Mandarin audio
Tamil audio… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper.whisper_transcriptions.reazonspeech.all.wer_10.0.vectorizedwhisper-rirmega-bench
Whisper-RIR-Mega: Paired Clean↔Reverberant Speech Robustness Benchmark
Dataset Summary
Whisper-RIR-Mega is a benchmark dataset of paired clean and reverberant speech for evaluating ASR robustness to room acoustics. Each sample consists of:
audio_clean: Clean speech (LibriSpeech test-clean, 16 kHz)
audio_reverb: Same utterance convolved with one RIR from RIR-Mega (v2)
text_ref: Ground-truth transcript
RIR metadata: rir_id, RT60, DRR, C50, etc. when available
Technical… See the full description on the dataset page: https://huggingface.co/datasets/mandipgoswami/whisper-rirmega-bench.whisperkit-evals-dataset
WhisperKit Evals Dataset
Overview
The WhisperKit Evals Dataset is a comprehensive collection of our speech recognition evaluation results, specifically designed to benchmark the performance of WhisperKit models across various devices and operating systems. This dataset provides detailed insights into performance and quality metrics, and model behavior under different conditions.
Dataset Structure
The dataset is organized into JSON files, each representing a… See the full description on the dataset page: https://huggingface.co/datasets/argmaxinc/whisperkit-evals-dataset.stt-pseudo-labeled-whisper-large-v3-multilingualThis collection includes over 189,000 hours of speech-to-text data in seven languages: English, French, Spanish, Portuguese, Italian, German, and Dutch
All segments were initially sorted by their IDs (timestamps). Adjacent segments from the same source were concatenated into 30-second chunks before being decoded using Whisper-Large-V3. The only exception was Common Voice, where segments were decoded individually before concatenation.
In total, over 288,000 hours of audio data were collected… See the full description on the dataset page: https://huggingface.co/datasets/bofenghuang/stt-pseudo-labeled-whisper-large-v3-multilingual.pseudolabel-malaysian-youtube-whisper-large-v3
Pseudolabel Malaysian Youtube videos using Whisper Large V3
Original dataset at https://huggingface.co/datasets/malaysia-ai/crawl-youtube, distributed pseudolabelled using 4x A100s
script at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text-semisupervised/pseudolabel-whisper
Each audio is 30 seconds.
Each audio saved in 16k sample rate.
whisper_transcriptions.reazon_speech_all.wer_10.0pseudolabel-malaysian-youtube-whisper-large-v3-timestamp
Pseudolabel Malaysian Youtube using Whisper Large V3 including Timestamp
how to prepare the dataset
wget https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp/resolve/main/prepared-pseudolabel.jsonl
huggingface-cli download --repo-type dataset \
--include 'output-audio-*.zip' \
--local-dir './' \
--max-workers 20 \
mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp
wget… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp.whisper_transcriptions.mlsmosaic-whisper-combinedAudio files from these links
https://huggingface.co/datasets/mesolitica/pseudolabel-malaya-speech-stt-train-whisper-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-imda-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-indonesian-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-nusantara-large-v3-timestamp
pseudolabel-dialects-youtube-whisper-large-v3
malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3
Pseudolabel malaysia-ai/malaysian-dialects-youtube using openai/whisper-large-v3
How to prepare the dataset
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \
malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3
wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3.german-asr-mixed-whisper
Dataset Card
Dataset Sources and Licensing
This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use.
Dataset Name
Original Source / Author
Link
TUDA-De German Speech Corpus
LT Group at UHH / TU Darmstadt
https://huggingface.co/datasets/uhhlt/Tuda-De
Mozilla Common Voice
Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/fosple/german-asr-mixed-whisper.
