datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anv_data_ke_kikuyu_scriptedcommon-voice-scripted-speech-kab-26-huge
Common Voice Scripted Speech 26.0 - Kabyle (Huge, Cleaned)
Full cleaned dataset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR. No speaker cap, no splits — all validated, cleaned, GlotLID-filtered clips.
Source
Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12)
Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective)
License: CC0-1.0
Generated: 2026-07-12
Cleaning Pipeline
Quality filter: ≥2 upvotes, 0 downvotes… See the full description on the dataset page: https://huggingface.co/datasets/boffire/common-voice-scripted-speech-kab-26-huge.common-voice-scripted-speech-quebec
common-voice-scripted-speech-quebec
This dataset is a filtered subset of the Mozilla Common Voice Scripted Speech 25.0 - French dataset. It exclusively contains audio clips from speakers with Canadian and Québécois accents.
Dataset Summary
Language: French (fr)
Total Clips: 25,198
Total Duration: 36.83 hours (132,578.64 seconds)
License: CC-0
Filtering Criteria
This subset was generated by extracting rows from the original cv-corpus-25.0-2026-03-09 dataset… See the full description on the dataset page: https://huggingface.co/datasets/thomasgauthier/common-voice-scripted-speech-quebec.cv-en-scripted-test-500
Common Voice English Scripted Test Set — 500 clips
n = 500 utterances · private eval set for ASR benchmarking
Source
Derived from Mozilla Common Voice Scripted Speech 25.0 — English (test split), downloaded via the Mozilla Data Collective API (dataset ID cmndapwry02jnmh07dyo46mot, 94 GB tarball).
Construction
Starting from the full CV 25.0 English test split (16,398 rows), a stratified 500-clip subset was produced using the same recipe as… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/cv-en-scripted-test-500.anv_scripted_multilingual-1common-voice-scripted-speech-kab-26-tiny
Common Voice Scripted Speech 26.0 - Kabyle (Cleaned)
This is a cleaned, speaker-disjoint subset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR.
Source
Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12)
Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective)
License: CC0-1.0
Generated: 2026-07-12
Cleaning Pipeline
Step
Input
Output
Filter
Quality filter
609,940
573,073
≥2 upvotes, 0 downvotes
Character… See the full description on the dataset page: https://huggingface.co/datasets/boffire/common-voice-scripted-speech-kab-26-tiny.scripted-malay-daily-use-speech-corpus
scripted-malay-daily-use-speech-corpus
Mirror for https://magichub.com/datasets/malay-scripted-speech-corpus-daily-use-sentence/, license is Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License
scripted-malay-daily-use-speech-corpus-whisper-formatanv_data_ke_kikuyu_scriptedanv_data_ke_kikuyu_scripted
