datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wav2vec2_common_voice_accents_3common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.Custom_common_voice_dataset_using_RVC
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion
Custom common_voice_v11 corpus with a custom voice was was created using RVC(Retrieval-Based Voice Conversion)
The model underwent 200 epochs of training, utilizing a total of 1 hour of audio clips. The data was scraped from Youtube.
The audio in the custom generated dataset is of a YouTuber named
Ajay Pandey
Description
license: cc0-1.0
language:
- hi… See the full description on the dataset page: https://huggingface.co/datasets/Aniket-Tathe-08/Custom_common_voice_dataset_using_RVC.common_voice_22_0
Common Voice Corpus 22.0
Originally from https://huggingface.co/datasets/fsicoli/common_voice_22_0, we mirror using multiple zip files also trimmed the silents.
How to prepare the dataset
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \
malaysia-ai/common_voice_22_0
wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/common_voice_22_0.common_voice_17_0
Common Voice Corpus 17.0
Mirror for mozilla-foundation/common_voice_17_0, easy to download and extract instead audio in parquet files.
How to prepare the dataset
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \
malaysia-ai/common_voice_17_0
wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3 unzip.py
common_voice_26_0_de
Mozilla Common Voice 26.0 - German (IPA & Clean Validated Subset)
Repacking version of Common Voice 26.0 German officialy published by Mozilla Data Collective, following Hugging Face Parquet Shards standard, with feature for listening to audio directly on the Web Hub, and the addition of a data column for the IPA transcription of each sentence.
📊 Dataset parameters
Origin: Mozilla Common Voice 26.0 (version 18/06/2026).
Data amount (Validated): 950,877 MP3 audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/common_voice_26_0_de.hausa_common_voiceThis dataset is from the common voice corpus 7.0 using the Hausa dataset
malayalam_common_voice_benchmarkingaccented_common_voicemozilla-common-voice-converted-to-parquet-ptcommon_voice_tamil_english-labeled-Data-filtered-v4dv_finetune_common_voice_13common_voice_16_0_testdb_ckb_untaggedcommonvoice-mnmozilla-common-voice-23-bel-texts-exportcommon-voice-ro-taggedcommonvoicecommon_voice_17_0_top1_testdb_ckb_taggedcommon_voice_13_0_id_taggedcommonvoice22-sidon-xcodec2common_voice_13_0_id_male_tagscommon_voice_17_0_top10_testdb_ckb_untagged_dataspeech_testcommon-voicecommonvoice22_sidon_becommon_voice_17_0_top10_testdb_ckb_taggedcommon_voice_13_0_id_tagscommon_voice_pt_dataspeechcommon_voice_16_0_testdb_ckb_untagged_dataspeech_testmozilla_commonvoice_naijaHausa2_preprocessed_train_batch_1curation-backup-review_common_voice25_dev
Dataset Card for curation-backup-review_common_voice25_dev
This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Using this dataset with Argilla
To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code:
import argilla as rg
ds… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/curation-backup-review_common_voice25_dev.
