CoolFace
Datasetpublic

Trelis/cv-en-scripted-test-500

Common Voice English Scripted Test Set — 500 clips n = 500 utterances · private eval set for ASR benchmarking Source Derived from Mozilla Common Voice Scripted Speech 25.0 — English (test split), downloaded via the Mozilla Data Collective API (dataset ID cmndapwry02jnmh07dyo46mot, 94 GB tarball). Construction Starting from the full CV 25.0 English test split (16,398 rows), a stratified 500-clip subset was produced using the same recipe as… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/cv-en-scripted-test-500.

sourceHugging Facecc0-1.0updated 4mo agoView on Hugging Face
0likes26downloads
Dataset Card

Common Voice English Scripted Test Set — 500 clips

n = 500 utterances · private eval set for ASR benchmarking

Source

Derived from Mozilla Common Voice Scripted Speech 25.0 — English (test split), downloaded via the Mozilla Data Collective API (dataset ID cmndapwry02jnmh07dyo46mot, 94 GB tarball).

Construction

Starting from the full CV 25.0 English test split (16,398 rows), a stratified 500-clip subset was produced using the same recipe as Trelis/cv-hi-test-500:

StepRule
Quality filterup_votes ≥ 2, down_votes = 0
Duration filter0.5 s ≤ duration ≤ 25 s
Speaker capmax 4 clips per client_id
StratificationProportional allocation across duration buckets (0–3 s, 3–6 s, 6–10 s, 10–25 s)
Seed42

Build script: merge-bench-baselines/scripts/build_cv_en_scripted.py

Columns

ColumnTypeDescription
audioAudio bytes (MP3)Speech recording
transcriptionstringReference text
idstringCV clip path stem
client_idstringAnonymised speaker hash
durationfloatDuration in seconds
age / gender / accentsstringSpeaker metadata (may be empty)
up_votes / down_votesintCommunity quality votes

License

CC0 1.0 — Mozilla Common Voice audio and text are released into the public domain by contributors.