datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Whisper-Hallucination
Whisper Hallucination and Repetition Probes
This is a benchmark. Every evaluation config is test — do not fine-tune on it.
lexicon_synth is the exception: synthetic training material with its own train/test
split, and not one of the eight benchmark arms.
To build training data, exclude the items in
benchmark/exclusions.json
(546 FMA tracks, 1,168 FSD50K ids, 2,620 LibriSpeech utterances, the Malay stems). The
benchmark draws FSD50K eval and FMA shards 0–1, so training can use… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Whisper-Hallucination.Whispered
Time-aligned multilingual ASR enrichment over Common Voice 17
A time-aligned, quality-scored enrichment layer over Common Voice 17
for 11 languages across 7 writing systems. Each row is one Common Voice clip with:
the human transcript (the ground-truth target, used as-is),
a whisper-large-v3 machine transcript (enrichment / agreement signal — not a replacement),
word- and segment-level timestamps from MMS forced alignment of the human transcript,
language-ID, WER/CER agreement… See the full description on the dataset page: https://huggingface.co/datasets/burakaydinofficial/Whispered.preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2605
5bfe2d098c8486d97fac8be76d86ec9146435245
train
56:46:32
50,557
589,095
11.7
31.9
techiaith/corpws-clllc-wlga
5d00294c31c78b1d7937bb2c2bc6cc70bc18d410
clips
48:20:49
27,579… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607.pseudolabel-malaya-speech-stt-train-whisper-large-v3apple-speechanalyzer-vs-whisper-cpp-mac
Apple SpeechAnalyzer vs whisper.cpp on Mac
Four complete speech-recognition benchmark runs over the same deterministic
40-speaker LibriSpeech test-clean snapshot:
Engine
Model path
WER
CER
Repeated median post-speech latency
Repeated p95
Apple SpeechAnalyzer
progressiveTranscription on macOS 26.5
1.98%
1.02%
125–132 ms
194–201 ms
whisper.cpp server
1.8.4 · ggml-small.en
4.28%
1.79%
122–125 ms
152–161 ms
Every run completed 40/40 clips with no failures. Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/apple-speechanalyzer-vs-whisper-cpp-mac.preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2601
main
train
52:38:17
48,608
573,772
11.8
27.7
techiaith/corpws-clllc-wlga
main
clips
20:00:52
18,905
228,962
12.1
10.5
techiaith/commonvoice_23_0_cy
main… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601.preprocessed-whisper-btb-cv-cvad-wlga-ca-2603
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2603
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2602
main
train
52:45:22
48,569
556,542
11.5
28.2
techiaith/corpws-clllc-wlga
2603_rc3
clips
41:54:21
29,446
447,780
15.2
22.4
techiaith/commonvoice_23_0_cy
main… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2603.preprocessed-whisper-btb-cv-cvad-wlga-ca-2606
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2606
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2605
main
train
57:14:31
50,934
594,564
11.7
29.5
techiaith/corpws-clllc-wlga
main
clips
52:27:50
29,794
544,066
18.3
27.1
techiaith/commonvoice_23_0_cy
main… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2606.preprocessed-whisper-btb-cv-cvad-wlga-ca-2602
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2602
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2602
main
train
52:45:22
48,569
556,542
11.5
32.2
techiaith/corpws-clllc-wlga
main
clips
20:02:39
18,851
228,230
12.1
12.2
techiaith/commonvoice_23_0_cy
main
train+dev+other_with_excluded… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2602.
