CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01willcai /wav2vec2_common_voice_accents_3tabular100K<n<1M0 likes2.2k downloads5y agoHugging Face02Peacockery /common-voice-scripted-speech-26 Common Voice Scripted Speech A row-normalized multilingual ASR dataset built from Mozilla Data Collective Common Voice Scripted Speech. Each upstream archive is converted to appendable parquet shards under data/<upstream_split>/, one shard per source archive and split, with audio bytes embedded in an audio struct column. Status Manifest languages: 60 Languages uploaded: 18 Columns audio (bytes, path) sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.tabularautomatic-speech-recognition100K<n<1M0 likes1.9k downloads3mo agoHugging Face03Aniket-Tathe-08 /Custom_common_voice_dataset_using_RVC Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion Custom common_voice_v11 corpus with a custom voice was was created using RVC(Retrieval-Based Voice Conversion) The model underwent 200 epochs of training, utilizing a total of 1 hour of audio clips. The data was scraped from Youtube. The audio in the custom generated dataset is of a YouTuber named Ajay Pandey Description license: cc0-1.0 language: - hi… See the full description on the dataset page: https://huggingface.co/datasets/Aniket-Tathe-08/Custom_common_voice_dataset_using_RVC.tabular10K<n<100K0 likes801 downloads3y agoHugging Face04malaysia-ai /common_voice_22_0 Common Voice Corpus 22.0 Originally from https://huggingface.co/datasets/fsicoli/common_voice_22_0, we mirror using multiple zip files also trimmed the silents. How to prepare the dataset huggingface-cli download --repo-type dataset \ --include '*.zip' \ --local-dir './' \ --max-workers 20 \ malaysia-ai/common_voice_22_0 wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py python3… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/common_voice_22_0.audio10M<n<100M1 likes543 downloads1y agoHugging Face05malaysia-ai /common_voice_17_0 Common Voice Corpus 17.0 Mirror for mozilla-foundation/common_voice_17_0, easy to download and extract instead audio in parquet files. How to prepare the dataset huggingface-cli download --repo-type dataset \ --include '*.zip' \ --local-dir './' \ --max-workers 20 \ malaysia-ai/common_voice_17_0 wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py python3 unzip.py audio1M<n<10M0 likes479 downloads1y agoHugging Face06q1805 /common_voice_26_0_de Mozilla Common Voice 26.0 - German (IPA & Clean Validated Subset) Repacking version of Common Voice 26.0 German officialy published by Mozilla Data Collective, following Hugging Face Parquet Shards standard, with feature for listening to audio directly on the Web Hub, and the addition of a data column for the IPA transcription of each sentence. 📊 Dataset parameters Origin: Mozilla Common Voice 26.0 (version 18/06/2026). Data amount (Validated): 950,877 MP3 audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/common_voice_26_0_de.tabularautomatic-speech-recognition100K<n<1M0 likes202 downloads1mo agoHugging Face07Arnold /hausa_common_voiceThis dataset is from the common voice corpus 7.0 using the Hausa dataset tabular1K<n<10K2 likes149 downloads5y agoHugging Face08kurianbenoy /malayalam_common_voice_benchmarkingtabular1K<n<10K1 likes112 downloads3y agoHugging Face09WillHeld /accented_common_voicetabular1K<n<10K1 likes66 downloads3y agoHugging Face10MatheusMarquesEiras /mozilla-common-voice-converted-to-parquet-pttabular10K<n<100K0 likes59 downloads11mo agoHugging Face11Lingalingeswaran /common_voice_tamil_english-labeled-Data-filtered-v4tabularaudio-classification1K<n<10K0 likes55 downloads2y agoHugging Face12Raziullah /dv_finetune_common_voice_13tabular1K<n<10K0 likes53 downloads3y agoHugging Face13roshna-omer /common_voice_16_0_testdb_ckb_untaggedtabular10K<n<100K0 likes43 downloads2y agoHugging Face14dgduksict /commonvoice-mnaudio1K<n<10K0 likes43 downloads11mo agoHugging Face15alex73 /mozilla-common-voice-23-bel-texts-exporttabulartext-generation100K<n<1M0 likes40 downloads10mo agoHugging Face16duckiduck /common-voice-ro-taggedtabular10K<n<100K0 likes39 downloads2y agoHugging Face17Amirjab21 /commonvoicetabular100K<n<1M0 likes37 downloads2y agoHugging Face18roshna-omer /common_voice_17_0_top1_testdb_ckb_taggedtabularn<1K0 likes29 downloads2y agoHugging Face19boringtaskai /common_voice_13_0_id_taggedtabular10K<n<100K0 likes26 downloads2y agoHugging Face20naive-puzzle /commonvoice22-sidon-xcodec2tabular100K<n<1M0 likes26 downloads10mo agoHugging Face21boringtaskai /common_voice_13_0_id_male_tagstabular1K<n<10K0 likes25 downloads2y agoHugging Face22roshna-omer /common_voice_17_0_top10_testdb_ckb_untagged_dataspeech_testtabular1K<n<10K0 likes23 downloads2y agoHugging Face23CS-224s /common-voicetabular10K<n<100K0 likes22 downloads2y agoHugging Face24archivartaunik /commonvoice22_sidon_betabular100K<n<1M0 likes22 downloads6mo agoHugging Face25roshna-omer /common_voice_17_0_top10_testdb_ckb_taggedtabular1K<n<10K0 likes20 downloads2y agoHugging Face26boringtaskai /common_voice_13_0_id_tagstabular10K<n<100K0 likes18 downloads2y agoHugging Face27lgris /common_voice_pt_dataspeechtabular100K<n<1M0 likes18 downloads2y agoHugging Face28roshna-omer /common_voice_16_0_testdb_ckb_untagged_dataspeech_testtabular10K<n<100K0 likes18 downloads2y agoHugging Face29EYEDOL /mozilla_commonvoice_naijaHausa2_preprocessed_train_batch_1tabular10K<n<100K0 likes18 downloads1y agoHugging Face30Reza2kn /curation-backup-review_common_voice25_dev Dataset Card for curation-backup-review_common_voice25_dev This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Using this dataset with Argilla To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code: import argilla as rg ds… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/curation-backup-review_common_voice25_dev.tabularn<1K1 likes17 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.