CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fsicoli /common_voice_15_0 Dataset Card for Common Voice Corpus 15.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 15. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_15_0.automatic-speech-recognition100B<n<1T6 likes48k downloads3y agoHugging Face02fsicoli /common_voice_22_0 Dataset Card for Common Voice Corpus 22.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_22_0.automatic-speech-recognition100B<n<1T20 likes28k downloads1y agoHugging Face03fsicoli /common_voice_17_0 Dataset Card for Common Voice Corpus 17.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 17. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_17_0.automatic-speech-recognition100B<n<1T19 likes16k downloads2y agoHugging Face04legacy-datasets /common_voiceCommon Voice is Mozilla's initiative to help teach machines how real people speak. The dataset currently consists of 7,335 validated hours of speech in 60 languages, but we’re always adding more voices and languages.automatic-speech-recognition100K<n<1M148 likes15k downloads2y agoHugging Face05fsicoli /common_voice_16_0 Dataset Card for Common Voice Corpus 16.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 16. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_16_0.automatic-speech-recognition100B<n<1T4 likes13k downloads3y agoHugging Face06mozilla-foundation /common_voice_17_0Effective October 2025, Mozilla Common Voice datasets are now exclusively available through Mozilla Data Collective. You can learn more about this change here. automatic-speech-recognition49 likes3.5k downloads11mo agoHugging Face07fsicoli /common_voice_19_0 Dataset Card for Common Voice Corpus 19.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 19. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_19_0.automatic-speech-recognition100B<n<1T9 likes2.8k downloads2y agoHugging Face08hezarai /common-voice-13-faThe Persian portion of the original CommonVoice 13 dataset at https://huggingface.co/datasets/mozilla-foundation/common_voice_13_0 Load # Using HF Datasets from datasets import load_dataset dataset = load_dataset("hezarai/common-voice-13-fa", split="train") # Using Hezar from hezar.data import Dataset dataset = Dataset.load("hezarai/common-voice-13-fa", split="train") audioautomatic-speech-recognition10K<n<100K1 likes2.7k downloads2y agoHugging Face09fsicoli /common_voice_21_0 Dataset Card for Common Voice Corpus 21.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 21. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_21_0.automatic-speech-recognition100B<n<1T2 likes2.2k downloads1y agoHugging Face10Theafricatechguy /common_voice_21_0 Dataset Card for Common Voice Corpus 21.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 21. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash… See the full description on the dataset page: https://huggingface.co/datasets/Theafricatechguy/common_voice_21_0.automatic-speech-recognition100B<n<1T0 likes2k downloads29d agoHugging Face11Peacockery /common-voice-scripted-speech-26 Common Voice Scripted Speech A row-normalized multilingual ASR dataset built from Mozilla Data Collective Common Voice Scripted Speech. Each upstream archive is converted to appendable parquet shards under data/<upstream_split>/, one shard per source archive and split, with audio bytes embedded in an audio struct column. Status Manifest languages: 60 Languages uploaded: 18 Columns audio (bytes, path) sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.tabularautomatic-speech-recognition100K<n<1M0 likes1.9k downloads3mo agoHugging Face12TTS-AGI /commonvoice22-sidon-dacvae CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents Source sarulab-speech/commonvoice22_sidon Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.audioautomatic-speech-recognition1M<n<10M1 likes1.3k downloads6mo agoHugging Face130x3 /common_voice_17_0 Dataset Card for Common Voice Corpus 17.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 17. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash… See the full description on the dataset page: https://huggingface.co/datasets/0x3/common_voice_17_0.automatic-speech-recognition100B<n<1T0 likes1.2k downloads4mo agoHugging Face14distil-whisper /common_voice_13_0-timestamped Distil Whisper: Common Voice 13 With Timestamps This is a variant of the Common Voice 13 dataset, augmented to return the pseudo-labelled Whisper Transcriptions alongside the original dataset elements. The pseudo-labelled transcriptions were generated by labelling the input audio data with the Whisper large-v2 model with greedy sampling and timestamp prediction. For information on how the original dataset was curated, refer to the original dataset card. Standalone Usage… See the full description on the dataset page: https://huggingface.co/datasets/distil-whisper/common_voice_13_0-timestamped.automatic-speech-recognition0 likes893 downloads3y agoHugging Face15quinnlue /common-voice-17-regmix-webdataset Common Voice 17 RegMix WebDataset Public, training-oriented WebDataset conversion of fsicoli/common_voice_17_0, pinned to source revision 8262c16bf297c87a9cd88c51997c4758ed7a8ba2. Layout Each language is a RegMix cluster at data/<language>/*.tar. Every sample is a pair with the same key: <key>.opus: mono Ogg Opus audio at 24 kbps, preserving the source sample rate (normally 48 kHz) <key>.json: UTF-8 training metadata and the complete original TSV row The source… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/common-voice-17-regmix-webdataset.automatic-speech-recognition10M<n<100M0 likes826 downloads25d agoHugging Face16Snagy22000 /common_voice_17_0 Dataset Card for Common Voice Corpus 17.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 17. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/Snagy22000/common_voice_17_0.automatic-speech-recognition100B<n<1T0 likes647 downloads6mo agoHugging Face17JacobLinCool /common_voice_19_0_zh-TW Common Voice Corpus 19.0 Chinese (Taiwan) The test set is the same as the original test set, while validated_without_test includes all validated examples except those with sentence IDs that appear in the test set. validated_without_test has about 50,000 examples in total, equivalent to approximately 44 hours, and is intended for use as the training set. test has about 5,000 examples, which is approximately 5 hours. audioautomatic-speech-recognition10K<n<100K3 likes615 downloads2y agoHugging Face18MohamedRashad /common-voice-18-arabic Dataset Card for Common Voice 18 – Arabic Edition Dataset Summary This dataset is an unofficial Arabic-only extraction of Mozilla Common Voice Corpus 18.0, prepared for Automatic Speech Recognition (ASR) research and development. It is derived from the original Common Voice 18 release and filtered to include Arabic (ar) speech data only, while preserving the original dataset structure, splits, and metadata fields. The dataset consists of validated, unvalidated, and… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/common-voice-18-arabic.audioautomatic-speech-recognition100K<n<1M5 likes475 downloads9mo agoHugging Face19Cnam-LMSSC /common_voice_13_french_phoneme Common Voice 13 French Phoneme Dataset Summary This dataset is a curated version of the French subset of Common Voice 13.0, enriched with a phonetic transcription column (phoneme). It was created by the Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) to support research in speech processing, specifically for tasks requiring phonetic alignment, phoneme recognition, and robust speech-to-text applications in French. The dataset retains the… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/common_voice_13_french_phoneme.audioautomatic-speech-recognition100K<n<1M1 likes382 downloads9mo agoHugging Face20xmodar /commonvoice-12.0-arabic-voice-converted Dataset Card for Voice Converted Arabic Common Voice 12.0 This dataset is derived from the Common Voice Arabic Corpus 12.0 and includes automatically diacritized transcriptions and phoneme representations for the original augmented audio data. The recordings feature Arabic text read aloud by users, where the text was initially undiacritized, allowing for potential reading errors. The diacritization and phonemes were generated automatically, resulting in a dataset that is valuable… See the full description on the dataset page: https://huggingface.co/datasets/xmodar/commonvoice-12.0-arabic-voice-converted.audioautomatic-speech-recognition100K<n<1M8 likes359 downloads2y agoHugging Face21Chillarmo /common_voice_20_armenian Common Voice 20 - Armenian This dataset is the Armenian portion of Mozilla's Common Voice 20.0 release, a massively multilingual collection of transcribed speech intended for speech technology research and development. Dataset Details Language: Armenian (hy) Source: Mozilla Common Voice Version: 20.0 License: CC0-1.0 audioautomatic-speech-recognition10K<n<100K1 likes351 downloads1y agoHugging Face22OpenFormosa /common_voice_25_zh-TW Common Voice Scripted Speech 25.0 - Chinese (Taiwan) This repository mirrors the Chinese (Taiwan) (zh-TW) portion of Mozilla Common Voice Scripted Speech 25.0 in Hugging Face datasets format. The audio has been embedded into Parquet and exposed as a Hugging Face Audio feature at 48 kHz. Official source page: Mozilla Data Collective - Common Voice Scripted Speech 25.0 - Chinese (Taiwan) Dataset Details Field Value Dataset ID cmn2g7eaj01fio10769r1m96n… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/common_voice_25_zh-TW.audioautomatic-speech-recognition100K<n<1M2 likes322 downloads4mo agoHugging Face23ysdede /commonvoice_17_tr_fixed Improving CommonVoice 17 Turkish Dataset I recently worked on enhancing the Mozilla CommonVoice 17 Turkish dataset to create a higher quality training set for speech recognition models.Here's an overview of my process and findings. Initial Analysis and Split Organization My first step was analyzing the dataset organization to understand its structure.Through analysis of filename stems as unique keys, I revealed and documented an important aspect of CommonVoice's design… See the full description on the dataset page: https://huggingface.co/datasets/ysdede/commonvoice_17_tr_fixed.audioautomatic-speech-recognition10K<n<100K10 likes297 downloads2y agoHugging Face24projecte-aina /commonvoice_benchmark_catalan_accentsA new presentation of the corpus Catalan Common Voice v17.0 - metadata annotated version with the splits redefined to benchmark ASR models with various Catalan accentautomatic-speech-recognition1M<n<10M3 likes278 downloads2y agoHugging Face25azeem-ahmed /Common_Voice_Corpus_22_0_Urdu Common Voice Corpus 22.0 - Urdu This dataset contains the Urdu subset of the Mozilla Common Voice 22.0 corpus, released in June 2025.It consists of crowdsourced speech recordings and their corresponding text transcriptions, collected to support open-source speech technology. Dataset Summary The Common Voice Corpus 22.0 Urdu dataset provides high-quality speech data for automatic speech recognition (ASR), speaker identification, and linguistic research in Urdu.It includes… See the full description on the dataset page: https://huggingface.co/datasets/azeem-ahmed/Common_Voice_Corpus_22_0_Urdu.audioautomatic-speech-recognition1 likes277 downloads1y agoHugging Face26vumichien /common_voice_large_jsut_jsss_css10 Dataset Card for vumichien/common_voice_large_jsut_jsss_css10 automatic-speech-recognition0 likes227 downloads4y agoHugging Face27ssahir /common_voice_13_0_dv_preprocessed Dataset Card for Common Voice Corpus 13.0 Dataset Summary The Common Voice dataset consists of a unique MP3 and corresponding text file. Many of the 27141 recorded hours in the dataset also include demographic metadata like age, sex, and accent that can help improve the accuracy of speech recognition engines. The dataset currently consists of 17689 validated hours in 108 languages, but more voices and languages are always added. Take a look at the Languages page to… See the full description on the dataset page: https://huggingface.co/datasets/ssahir/common_voice_13_0_dv_preprocessed.automatic-speech-recognition1K<n<10K0 likes224 downloads3y agoHugging Face28btsee /common-voice-26-mn Common Voice 26.0 Mongolian (cleaned) A quality-filtered, normalised subset of Mozilla Common Voice Corpus 26.0, Mongolian, prepared for training Mongolian (Khalkha Cyrillic) text-to-speech with oron-tts. Built by oron-cleaner. Every threshold was calibrated on this corpus, and every number and column on this page is read from the shipped data rather than asserted. from datasets import load_dataset ds = load_dataset("btsee/common-voice-26-mn", split="train") print(ds[0]["text"]… See the full description on the dataset page: https://huggingface.co/datasets/btsee/common-voice-26-mn.audiotext-to-speech10K<n<100K1 likes224 downloads20d agoHugging Face29echodict /common_voice_11_0 Dataset Card for Common Voice Corpus 11.0 Dataset Summary The Common Voice dataset consists of a unique MP3 and corresponding text file. Many of the 24210 recorded hours in the dataset also include demographic metadata like age, sex, and accent that can help improve the accuracy of speech recognition engines. The dataset currently consists of 16413 validated hours in 100 languages, but more voices and languages are always added. Take a look at the Languages page to… See the full description on the dataset page: https://huggingface.co/datasets/echodict/common_voice_11_0.automatic-speech-recognition1M<n<10M2 likes215 downloads5mo agoHugging Face30q1805 /common_voice_26_0_de Mozilla Common Voice 26.0 - German (IPA & Clean Validated Subset) Repacking version of Common Voice 26.0 German officialy published by Mozilla Data Collective, following Hugging Face Parquet Shards standard, with feature for listening to audio directly on the Web Hub, and the addition of a data column for the IPA transcription of each sentence. 📊 Dataset parameters Origin: Mozilla Common Voice 26.0 (version 18/06/2026). Data amount (Validated): 950,877 MP3 audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/common_voice_26_0_de.tabularautomatic-speech-recognition100K<n<1M0 likes202 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.