datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cv_corpus_v22
Dataset Card for Common Voice Corpus 22.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
NOTE: currently converting to parquet for convenience.. WIP
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/mort666/cv_corpus_v22.cv-v1.0-segment
CommonVoice v1 Phone-Segment Alignments
Phone-level time alignments for 10 languages of Mozilla Common Voice,
packaged in a canonical segmentation schema with embedded 16 kHz audio. The
phone boundaries come from the charsiu/cv_ali
release of MFA alignments; the audio and transcripts come from
Common Voice Corpus 13.0 (2023-03-09).
Dataset summary
lang
train rows
train hrs
val rows
val hrs
test rows
test hrs
en
1,008,669
1,354.0
3,537
4.9
1,285
1.7
rw… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/cv-v1.0-segment.cv-22-deGerman split of Common Voice 22. cc0 license
cv-corpus-17.0-zh-TW-client_id-grouped
cv-corpus-17.0-zh-TW-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-zh-TW-client_id-grouped.cv-corpus-1.0-en-client_id-grouped
cv-corpus-1.0-en-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 60 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-1.0-en-client_id-grouped.nb-NO-asr-cv
Norwegian Bokmål ASR (Common Voice 22, filtered + rebalanced)
Norwegian Bokmål (nb-NO) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via
the open NbAiLab/NPSC mirror. Built to fine-tune
streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b).
Splits
Split
Clips
Hours
train
44302
86.9
dev
453
0.9
test
1211
2.2
train = the official validated train split + the filtered other bucket + the excess
dev/test speakers: Common… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/nb-NO-asr-cv.cv-corpus-17.0-zh-CN-client_id-grouped
cv-corpus-17.0-zh-CN-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-zh-CN-client_id-grouped.cv22-opus
Common Voice for 🇺🇦 Ukrainian (OPUS)
Ukrainian validated subset of Common Voice 22
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: 89248
Total duration: 115h 5m 9s
nl-asr-cv
Dutch ASR (Common Voice, speaker-disjoint splits)
Dutch (nl) speech for ASR, built from Mozilla Common Voice (CC0) via the
open fsicoli/common_voice_17_0 mirror. Built to fine-tune tiny
ASR models (e.g. openai/whisper-tiny).
Splits
Split
Hours
train
94.4
dev
0.8
test
2.0
Held-out dev/test are disjoint from train by both speaker and sentence. Common Voice's
official dev/test are capped by whole speakers (dev ~0.75h, test
~2.0h) with the… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/nl-asr-cv.cv-corpus-17.0-ja-client_id-grouped
cv-corpus-17.0-ja-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-ja-client_id-grouped.cv-corpus-25.0-ja
Mozilla Common Voice 25.0 - Japanese Test Set (Complete)
Dataset Description
Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace.
Key Features
Size: 9,019 validated test utterances
Coverage: 100% of official Common Voice 25.0 Japanese test split
Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.cy-asr-cv
Welsh ASR (Common Voice 22, filtered + rebalanced)
Welsh (cy) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via
the open fsicoli/common_voice_22_0 mirror. Built to fine-tune
streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b).
Splits
Split
Clips
Hours
train
?
50.1
dev
?
0.8
test
?
2.0
train = the official validated train split + the filtered other bucket + the excess
dev/test speakers: Common Voice's official… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/cy-asr-cv.frisian-asr-cv22
Frisian ASR (Common Voice 22, filtered)
Open Standard West Frisian (fy-NL) speech for ASR, built from Mozilla Common Voice 22.0
(CC0). The validated training split is augmented with the unvalidated other bucket, which is
auto-filtered by CTC agreement with the known prompt using a Frisian-specialized wav2vec2 model.
Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b).
Splits
Split
Clips
Hours
Composition
train
29,929… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/frisian-asr-cv22.preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2605
5bfe2d098c8486d97fac8be76d86ec9146435245
train
56:46:32
50,557
589,095
11.7
31.9
techiaith/corpws-clllc-wlga
5d00294c31c78b1d7937bb2c2bc6cc70bc18d410
clips
48:20:49
27,579… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607.cv17_es_other_automatically_verifiedSplit called -other- of the Spanish Common Voice v17.0 that was automatically verified
using various ASR system.sv-SE-asr-cv
Swedish ASR (Common Voice 22, filtered + rebalanced)
Swedish (sv-SE) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via
the open fsicoli/common_voice_22_0 mirror. Built to fine-tune
streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b).
Splits
Split
Clips
Hours
train
22166
26.1
dev
694
0.8
test
1602
2.0
train = the official validated train split + the filtered other bucket + the excess
dev/test speakers: Common Voice's… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/sv-SE-asr-cv.CV18-NER
CV-18 NER
CV-18 NER is the first publicly available dataset for Named Entity Recognition (NER) from Arabic speech. It was created by augmenting the Arabic Common Voice 18 corpus with manual NER annotations following the fine-grained Wojood schema, which covers 21 entity types.
The dataset provides a benchmark for evaluating both pipeline systems (ASR + text NER) and end-to-end speech NER models. It is particularly valuable for research in low-resource settings and morphologically… See the full description on the dataset page: https://huggingface.co/datasets/Elyadata/CV18-NER.da-asr-cv
Danish ASR (Common Voice 22, filtered + rebalanced)
Danish (da) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via
the open fsicoli/common_voice_22_0 mirror. Built to fine-tune
streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b).
Splits
Split
Clips
Hours
train
8270
10.0
dev
734
0.9
test
1593
2.1
train = the official validated train split + the filtered other bucket + the excess
dev/test speakers: Common Voice's… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/da-asr-cv.processed-cv-17-th-130k
processed-cv-17-th-130k
Cleaned Thai split of Mozilla Common Voice 17: 130,551 utterances (117,536 train / 3,950 dev / 9,065 test) with transcripts, ready for ASR training.
Format
Field
Description
sentence
Transcript in Thai
audio
Audio clip
Usage
from datasets import load_dataset
ds = load_dataset("Porameht/processed-cv-17-th-130k")
Source and license
Derived from Mozilla Common Voice 17.0 (Thai), released under… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/processed-cv-17-th-130k.spite-CV16-TP9B
Spite Dataset
Pseudolabeled speech translation data with quality annotations from multiple metrics. This version uses transcripts from Common Voice 16.1 and translations from Tower-Plus-9B.
Configs
en_de
en_es
en_fr
en_it
en_ko
en_nl
en_pt
en_ru
en_zh
Usage
from datasets import load_dataset
ds = load_dataset("bpop/spite-CV16-Euro9B", "en_pt")
cv10-uk-testset-clean
The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦
Overview
This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios.
All audios have been checked by a human to be sure that they are correct.
This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk
Community
Discord: https://bit.ly/discord-uds
Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.cv-mn-24.0
cv-mn-24.0
Mongolian subset of the Mozilla Common Voice speech recognition dataset.
Dataset Statistics
Total samples: 6,018Total duration: 9h 7m 2s (9.12 h)
Per-split breakdown
Split
Samples
Total Duration
Avg Duration
train
2,188
3h 7m 34s (3.13 h)
5.14 s
validation
1,896
2h 54m 41s (2.91 h)
5.53 s
test
1,934
3h 4m 47s (3.08 h)
5.73 s
cv-tr-eval
Common Voice Turkish Eval
4,825 Turkish test clips (~45 MB, 16 kHz) in the Mozilla Common Voice schema:
transcription, duration, up_votes / down_votes, and the age / gender
/ accent speaker attributes. Schema in the YAML header above.
An evaluation-only Turkish counterpart to
lahaja-eval;
never trained on. Used to sanity-check Turkish ASR quality on real human
speech, which matters here because the v0.3 training corpus is entirely
synthetic TTS and the released model is… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/cv-tr-eval.spite-CV16-Euro9B
Spite Dataset
Pseudolabeled speech translation data with quality annotations from multiple metrics. This version uses transcripts from Common Voice 16.1 and translations from EuroLLM-9B-Instruct.
Configs
en_de
en_es
en_fr
en_it
en_ko
en_nl
en_pt
en_ru
en_zh
Usage
from datasets import load_dataset
ds = load_dataset("bpop/spite-CV16-Euro9B", "en_pt")
cv-en-scripted-test-500
Common Voice English Scripted Test Set — 500 clips
n = 500 utterances · private eval set for ASR benchmarking
Source
Derived from Mozilla Common Voice Scripted Speech 25.0 — English (test split), downloaded via the Mozilla Data Collective API (dataset ID cmndapwry02jnmh07dyo46mot, 94 GB tarball).
Construction
Starting from the full CV 25.0 English test split (16,398 rows), a stratified 500-clip subset was produced using the same recipe as… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/cv-en-scripted-test-500.stt-calibration
STT Calibration Dataset
Tiny calibration dataset for PersonalAssistant STT service. Used on first run to auto-tune speculative pre-transcription parameters.
Contents
File
Duration
Size
Purpose
short.wav
3.5s
110KB
RTF measurement + VAD onset latency
long.wav
23.3s
729KB
Split quality calibration (whole vs split comparison)
very_long.wav
56.8s
1.8MB
Multi-split calibration (find minimum safe split interval)
manifest.json
-
2KB
Sample metadata + reference… See the full description on the dataset page: https://huggingface.co/datasets/cvxhull/stt-calibration.cv17_sw_kenyan_sample
Common Voice 17.0 — Swahili (Kenyan Sample)
This dataset is a filtered sample of the Mozilla Common Voice 17.0 corpus, focusing on Swahili (sw) speech with Kenyan voices.
It has been subsetted for experimentation and prototyping in ASR (Automatic Speech Recognition) models targeting speech-impaired users in Kenya, covering Kenyan English and Kiswahili.
Dataset Summary
Language: Kiswahili (Swahili, sw)
Accent/Region: Kenyan speakers
Domain: Conversational… See the full description on the dataset page: https://huggingface.co/datasets/Veronica1NW/cv17_sw_kenyan_sample.cv_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/cv_chunked
Aligned dataset: instinct-org/cv_chunked_nfa_aligned
Rows: 71097 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans
nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_nfa_aligned.cv_chunked
cv_chunked
This is a gated Uzbek chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked.cv_chunked_speech_restorised
cv_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_speech_restorised.
