datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common_voice_17_0common_voice_15_0
Dataset Card for Common Voice Corpus 15.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 15. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_15_0.common_voice_22_0
Dataset Card for Common Voice Corpus 22.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_22_0.common_voiceCommon Voice is Mozilla's initiative to help teach machines how real people speak.
The dataset currently consists of 7,335 validated hours of speech in 60 languages, but we’re always adding more voices and languages.common_voice_17_0
Dataset Card for Common Voice Corpus 17.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 17. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_17_0.common_voice_16_0
Dataset Card for Common Voice Corpus 16.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 16. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_16_0.common_voice_21_0uyghur-common-voice-tts
Uyghur Common Voice TTS Dataset
A cleaned and processed Text-to-Speech (TTS) dataset for the Uyghur language, derived from Mozilla Common Voice.
Dataset Summary
Property
Value
Language
Uyghur (ug)
Total Samples
43,054
Train Samples
40,901
Validation Samples
2,153
Audio Format
WAV
Source
Mozilla Common Voice
License
CC0-1.0
Dataset Structure
/
├── train.jsonl # Training data (40,901 samples)
├── val.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-common-voice-tts.common_voice_speak_textcommon_voice_17_0Effective October 2025, Mozilla Common Voice datasets are now exclusively available through Mozilla Data Collective. You can learn more about this change here.
common_voice_16_0common-voice-13-faThe Persian portion of the original CommonVoice 13 dataset at https://huggingface.co/datasets/mozilla-foundation/common_voice_13_0
Load
# Using HF Datasets
from datasets import load_dataset
dataset = load_dataset("hezarai/common-voice-13-fa", split="train")
# Using Hezar
from hezar.data import Dataset
dataset = Dataset.load("hezarai/common-voice-13-fa", split="train")
common_voice_19_0
Dataset Card for Common Voice Corpus 19.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 19. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_19_0.common_voice_21_0_minicommon-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.common_voice_21_0
Dataset Card for Common Voice Corpus 21.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 21. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash… See the full description on the dataset page: https://huggingface.co/datasets/Theafricatechguy/common_voice_21_0.wav2vec2_common_voice_accents_3commonvoice22_sidon
CV22-Sidon
Overview
This dataset hosts a release of Mozilla Common Voice 22 restored with the Sidon speech restoration model.
Source: Mozilla Common Voice 22.0
Processing: Sidon denoising (sarulab-speech/sidon-v0.1) with 21 s chunks and 48 kHz reconstruction
Format: WebDataset shards (.tar.gz)
Manifest: paths.yaml enumerates every shard path for Hugging Face–style loading
License: Original Common Voice license (CC0 1.0)
Languages
137 language folders are… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon.common_voice_21_0
Dataset Card for Common Voice Corpus 21.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 21. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_21_0.Custom_Common_Voice_16.0_dataset_using_RVC_14min_data
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion
Custom common_voice_v16 corpus with a custom voice was was created using RVC(Retrieval-Based Voice Conversion)
The model underwent 200 epochs of training, utilizing a total of 14min of audio clips. The data was scraped from Youtube.
The audio in the custom generated dataset is of a YouTuber named
Ajay Pandey
Description
license: cc0-1.0
language:
- hi… See the full description on the dataset page: https://huggingface.co/datasets/Aniket-Tathe-08/Custom_Common_Voice_16.0_dataset_using_RVC_14min_data.common_voice_10_1_th_augmented_pitch
Dataset Card for "common_voice_10_1_th_augmented_pitch"
More Information needed
common_voice_zh_hk_processed
Dataset Card for "common_voice_zh_hk_processed"
More Information needed
common_voice_17_0
Dataset Card for Common Voice Corpus 17.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 17. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/0x3/common_voice_17_0.CommonVoice-POSTPROCESS-f0ecfc0ccommonvoice22-sidon-dacvae
CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents
Source
sarulab-speech/commonvoice22_sidon
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.VIVOS_CommonVoice_FOSD_Control_processed_dataset
Dataset Card for "VIVOS_CommonVoice_FOSD_Control_processed_dataset"
More Information needed
CommonVoice22-Sidon-HQ
CommonVoice22-Sidon-HQ
The top-DNSMOS slice of sarulab-speech/commonvoice22_sidon —
Mozilla Common Voice 22.0 restored to 48 kHz by sarulab-speech/sidon-v0.1,
then scored clip-by-clip with DNSMOS P.835 and cut down to only the cleanest utterances.
All 14,946,932 source clips (20,746 h, 2.61 TB of FLAC) were scored;
613,305 (4.1%, 1,021 h) passed and are published here.
Filter
DNSMOS P.835 (speechmos, ONNX) on a single centred 10 s window at 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/CommonVoice22-Sidon-HQ.common_voice_13_0-timestamped
Distil Whisper: Common Voice 13 With Timestamps
This is a variant of the Common Voice 13 dataset, augmented to return the pseudo-labelled Whisper
Transcriptions alongside the original dataset elements. The pseudo-labelled transcriptions were generated by
labelling the input audio data with the Whisper large-v2
model with greedy sampling and timestamp prediction. For information on how the original dataset was curated, refer to the original
dataset card.
Standalone Usage… See the full description on the dataset page: https://huggingface.co/datasets/distil-whisper/common_voice_13_0-timestamped.common-voice-17-regmix-webdataset
Common Voice 17 RegMix WebDataset
Public, training-oriented WebDataset conversion of fsicoli/common_voice_17_0, pinned to source revision 8262c16bf297c87a9cd88c51997c4758ed7a8ba2.
Layout
Each language is a RegMix cluster at data/<language>/*.tar. Every sample is a pair with the same key:
<key>.opus: mono Ogg Opus audio at 24 kbps, preserving the source sample rate (normally 48 kHz)
<key>.json: UTF-8 training metadata and the complete original TSV row
The source… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/common-voice-17-regmix-webdataset.common_voice_13_0Effective October 2025, Mozilla Common Voice datasets are now exclusively available through [Mozilla Data Collective](https://datacollective.mozillafoundation.org]. You can learn more about this change here.
