datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
whisper_transcriptions.reazon_speech_allwhisper_transcriptions.reazonspeech.alllibrispeech_asr-noise
Dataset Card for "librispeech_asr-noise"
More Information needed
whisper_transcriptions.reazonspeech.all.wer_10.0earnings22
Dataset Card for Earnings 22
Dataset Summary
Earnings-22 provides a free-to-use benchmark of real-world, accented audio to bridge academic and industrial research.
This dataset contains 125 files totalling roughly 119 hours of English language earnings calls from global countries.
This dataset provides the full audios, transcripts, and accompanying metadata such as ticker symbol, headquarters country,
and our defined "Language Region".
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/distil-whisper/earnings22.meanwhile
Dataset Card for "meanwhile"
This dataset consists of 64 segments from The Late Show with Stephen Colbert. This dataset was published as
part of the Whisper release by OpenAI. See page 19 of the Whisper paper
for details.
linustechtips-transcript-audio
Dataset Card for "linustechtips"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips.
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.lex-fridman-podcast-transcript-audio
Dataset Card for "lexFridmanPodcast-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.
Data Fields
The dataset is composed by:
id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.Malaysian-STT-Whisper
Malaysian STT Whisper format
Heavy postprocessing and post-translation to improve pseudolabeled Whisper Large V3. Also include word level timestamp.
Postprocessing
Check repetitive trigrams.
Verify Voice Activity using Silero-VAD.
Verify scores using Force Alignment.
Post-translation
We use mesolitica/nanot5-base-malaysian-translation-v2.1.
Dataset involved
Malaysian context v2
Singaporean context
Indonesian context
Mandarin audio
Tamil audio… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper.pseudolabel-malaysian-youtube-whisper-large-v3-timestamp
Pseudolabel Malaysian Youtube using Whisper Large V3 including Timestamp
how to prepare the dataset
wget https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp/resolve/main/prepared-pseudolabel.jsonl
huggingface-cli download --repo-type dataset \
--include 'output-audio-*.zip' \
--local-dir './' \
--max-workers 20 \
mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp
wget… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp.whisper_transcriptions.mlsmosaic-whisper-combinedAudio files from these links
https://huggingface.co/datasets/mesolitica/pseudolabel-malaya-speech-stt-train-whisper-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-imda-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-indonesian-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-nusantara-large-v3-timestamp
pseudolabel-dialects-youtube-whisper-large-v3
malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3
Pseudolabel malaysia-ai/malaysian-dialects-youtube using openai/whisper-large-v3
How to prepare the dataset
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \
malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3
wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3.german-asr-mixed-whisper
Dataset Card
Dataset Sources and Licensing
This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use.
Dataset Name
Original Source / Author
Link
TUDA-De German Speech Corpus
LT Group at UHH / TU Darmstadt
https://huggingface.co/datasets/uhhlt/Tuda-De
Mozilla Common Voice
Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/fosple/german-asr-mixed-whisper.whisperkit-test-datawhisper_transcriptions.reazonspeech.large.wer_10.0instruction-convert-audio-whispervq-llama3.2-compressindic-superb-whisperlibrispeech_asr-prompted
Dataset Card for "librispeech_asr-prompted"
More Information needed
tedlium-prompted
Dataset Card for "tedlium-prompted"
More Information needed
whisper-dataset-ytb-uk
Dataset Card for Dataset Name
This dataset is collected from youtube.
knesset-plenums-whisper-training
Dataset Card for ivrit.ai - Knesset Plenums Whisper Training
This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset.
This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less.
Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription.
The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.Whisper-Hallucination
Whisper Hallucination and Repetition Probes
This is a BENCHMARK. Every evaluation config is test — do not fine-tune on it.
(The one exception is lexicon_synth, which is synthetic training material and ships its
own train/test split. It is not one of the eight benchmark arms — see below.)
Training on these clips invalidates every number you would then report. Build training
data separately from the same source corpora, excluding the items listed in
benchmark/exclusions.json in… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Whisper-Hallucination.whisperspeech
The WhisperSpeech Dataset
This dataset contains data to train SPEAR TTS-like text-to-speech models that utilized semantic tokens derived from the OpenAI Whisper
speech recognition model.
We currently provide semantic and acoustic tokens for the LibriLight and LibriTTS datasets (English only).
Acoustic tokens:
24kHz EnCodec 6kbps (8 quantizers)
Semantic tokens:
Whisper tiny VQ bottleneck trained on a subset of LibriLight
Available LibriLight subsets:
small/medium/large… See the full description on the dataset page: https://huggingface.co/datasets/collabora/whisperspeech.bambara-whisper-featuresMalaysian-STT-Whisper-Stage2
Malaysian STT Whisper Stage 2
Extra dataset to compliment mesolitica/Malaysian-STT-Whisper.
This dataset is stronger in confidence and suitable for second stage / annealing finetuning.
how to prepare the dataset
huggingface-cli download \
mesolitica/Malaysian-STT-Whisper-Stage2 \
--include "*.zip" \
--repo-type "dataset" \
--local-dir './'
huggingface-cli download \
mesolitica/Malaysian-Multiturn-Chat-Assistant \
--include "*.zip" \
--exclude "voice/*.zip" \
--repo-type… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper-Stage2.Whisper-fine-tune-2libri-proc-whisperinstruction-convert-audio-whispervq-llama3.2instruction-convert-audio-whispervq-llama3.2-dedup
