datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cv-en-scripted-test-500
Common Voice English Scripted Test Set — 500 clips
n = 500 utterances · private eval set for ASR benchmarking
Source
Derived from Mozilla Common Voice Scripted Speech 25.0 — English (test split), downloaded via the Mozilla Data Collective API (dataset ID cmndapwry02jnmh07dyo46mot, 94 GB tarball).
Construction
Starting from the full CV 25.0 English test split (16,398 rows), a stratified 500-clip subset was produced using the same recipe as… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/cv-en-scripted-test-500.preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2601
main
train
52:38:17
48,608
573,772
11.8
27.7
techiaith/corpws-clllc-wlga
main
clips
20:00:52
18,905
228,962
12.1
10.5
techiaith/commonvoice_23_0_cy
main… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601.
