CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mort666 /cv_corpus_v22 Dataset Card for Common Voice Corpus 22.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. NOTE: currently converting to parquet for convenience.. WIP Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/mort666/cv_corpus_v22.audioautomatic-speech-recognition1M<n<10M0 likes5.4k downloads9mo agoHugging Face02changelinglab /cv-v1.0-segment CommonVoice v1 Phone-Segment Alignments Phone-level time alignments for 10 languages of Mozilla Common Voice, packaged in a canonical segmentation schema with embedded 16 kHz audio. The phone boundaries come from the charsiu/cv_ali release of MFA alignments; the audio and transcripts come from Common Voice Corpus 13.0 (2023-03-09). Dataset summary lang train rows train hrs val rows val hrs test rows test hrs en 1,008,669 1,354.0 3,537 4.9 1,285 1.7 rw… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/cv-v1.0-segment.audioautomatic-speech-recognition1M<n<10M3 likes2k downloads5mo agoHugging Face03fidoriel /cv-22-deGerman split of Common Voice 22. cc0 license audioautomatic-speech-recognition100K<n<1M3 likes793 downloads1y agoHugging Face04masuidrive /cv-corpus-17.0-zh-TW-client_id-grouped cv-corpus-17.0-zh-TW-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-zh-TW-client_id-grouped.audioautomatic-speech-recognition10K<n<100K1 likes328 downloads2y agoHugging Face05masuidrive /cv-corpus-1.0-en-client_id-grouped cv-corpus-1.0-en-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 60 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-1.0-en-client_id-grouped.audioautomatic-speech-recognition100K<n<1M1 likes311 downloads2y agoHugging Face06LokaalHub /nb-NO-asr-cv Norwegian Bokmål ASR (Common Voice 22, filtered + rebalanced) Norwegian Bokmål (nb-NO) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via the open NbAiLab/NPSC mirror. Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b). Splits Split Clips Hours train 44302 86.9 dev 453 0.9 test 1211 2.2 train = the official validated train split + the filtered other bucket + the excess dev/test speakers: Common… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/nb-NO-asr-cv.audioautomatic-speech-recognition10K<n<100K0 likes308 downloads3mo agoHugging Face07masuidrive /cv-corpus-17.0-zh-CN-client_id-grouped cv-corpus-17.0-zh-CN-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-zh-CN-client_id-grouped.audioautomatic-speech-recognition100K<n<1M3 likes300 downloads2y agoHugging Face08speech-uk /cv22-opus Common Voice for 🇺🇦 Ukrainian (OPUS) Ukrainian validated subset of Common Voice 22 Community Discord: https://bit.ly/discord-uds Speech Recognition: https://t.me/speech_recognition_uk Speech Synthesis: https://t.me/speech_synthesis_uk Stats Total files processed: 89248 Total duration: 115h 5m 9s audioautomatic-speech-recognition10K<n<100K0 likes222 downloads6mo agoHugging Face09LokaalHub /nl-asr-cv Dutch ASR (Common Voice, speaker-disjoint splits) Dutch (nl) speech for ASR, built from Mozilla Common Voice (CC0) via the open fsicoli/common_voice_17_0 mirror. Built to fine-tune tiny ASR models (e.g. openai/whisper-tiny). Splits Split Hours train 94.4 dev 0.8 test 2.0 Held-out dev/test are disjoint from train by both speaker and sentence. Common Voice's official dev/test are capped by whole speakers (dev ~0.75h, test ~2.0h) with the… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/nl-asr-cv.audioautomatic-speech-recognition10K<n<100K0 likes212 downloads3mo agoHugging Face10masuidrive /cv-corpus-17.0-ja-client_id-grouped cv-corpus-17.0-ja-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-ja-client_id-grouped.audioautomatic-speech-recognition10K<n<100K2 likes187 downloads2y agoHugging Face11FluidInference /cv-corpus-25.0-ja Mozilla Common Voice 25.0 - Japanese Test Set (Complete) Dataset Description Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace. Key Features Size: 9,019 validated test utterances Coverage: 100% of official Common Voice 25.0 Japanese test split Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.audioautomatic-speech-recognition1K<n<10K0 likes170 downloads6mo agoHugging Face12LokaalHub /cy-asr-cv Welsh ASR (Common Voice 22, filtered + rebalanced) Welsh (cy) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via the open fsicoli/common_voice_22_0 mirror. Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b). Splits Split Clips Hours train ? 50.1 dev ? 0.8 test ? 2.0 train = the official validated train split + the filtered other bucket + the excess dev/test speakers: Common Voice's official… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/cy-asr-cv.audioautomatic-speech-recognition10K<n<100K0 likes135 downloads3mo agoHugging Face13LokaalHub /frisian-asr-cv22 Frisian ASR (Common Voice 22, filtered) Open Standard West Frisian (fy-NL) speech for ASR, built from Mozilla Common Voice 22.0 (CC0). The validated training split is augmented with the unvalidated other bucket, which is auto-filtered by CTC agreement with the known prompt using a Frisian-specialized wav2vec2 model. Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b). Splits Split Clips Hours Composition train 29,929… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/frisian-asr-cv22.audioautomatic-speech-recognition10K<n<100K0 likes105 downloads4mo agoHugging Face14DewiBrynJones /preprocessed-whisper-btb-cv-cvad-wlga-ca-2607 Dataset Card Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607 Revision: main Dataset Statistics Train Split Statistics Dataset Revision Split Duration (HH:MM:SS) Clips Words Words/Clip % DewiBrynJones/banc-trawsgrifiadau-bangor-2605 5bfe2d098c8486d97fac8be76d86ec9146435245 train 56:46:32 50,557 589,095 11.7 31.9 techiaith/corpws-clllc-wlga 5d00294c31c78b1d7937bb2c2bc6cc70bc18d410 clips 48:20:49 27,579… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607.tabularautomatic-speech-recognition100K<n<1M0 likes94 downloads1mo agoHugging Face15projecte-aina /cv17_es_other_automatically_verifiedSplit called -other- of the Spanish Common Voice v17.0 that was automatically verified using various ASR system.automatic-speech-recognition100K<n<1M2 likes91 downloads1y agoHugging Face16LokaalHub /sv-SE-asr-cv Swedish ASR (Common Voice 22, filtered + rebalanced) Swedish (sv-SE) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via the open fsicoli/common_voice_22_0 mirror. Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b). Splits Split Clips Hours train 22166 26.1 dev 694 0.8 test 1602 2.0 train = the official validated train split + the filtered other bucket + the excess dev/test speakers: Common Voice's… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/sv-SE-asr-cv.audioautomatic-speech-recognition10K<n<100K0 likes91 downloads3mo agoHugging Face17Elyadata /CV18-NER CV-18 NER CV-18 NER is the first publicly available dataset for Named Entity Recognition (NER) from Arabic speech. It was created by augmenting the Arabic Common Voice 18 corpus with manual NER annotations following the fine-grained Wojood schema, which covers 21 entity types. The dataset provides a benchmark for evaluating both pipeline systems (ASR + text NER) and end-to-end speech NER models. It is particularly valuable for research in low-resource settings and morphologically… See the full description on the dataset page: https://huggingface.co/datasets/Elyadata/CV18-NER.automatic-speech-recognition1 likes62 downloads4mo agoHugging Face18LokaalHub /da-asr-cv Danish ASR (Common Voice 22, filtered + rebalanced) Danish (da) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via the open fsicoli/common_voice_22_0 mirror. Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b). Splits Split Clips Hours train 8270 10.0 dev 734 0.9 test 1593 2.1 train = the official validated train split + the filtered other bucket + the excess dev/test speakers: Common Voice's… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/da-asr-cv.audioautomatic-speech-recognition10K<n<100K0 likes61 downloads3mo agoHugging Face19Porameht /processed-cv-17-th-130k processed-cv-17-th-130k Cleaned Thai split of Mozilla Common Voice 17: 130,551 utterances (117,536 train / 3,950 dev / 9,065 test) with transcripts, ready for ASR training. Format Field Description sentence Transcript in Thai audio Audio clip Usage from datasets import load_dataset ds = load_dataset("Porameht/processed-cv-17-th-130k") Source and license Derived from Mozilla Common Voice 17.0 (Thai), released under… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/processed-cv-17-th-130k.audioautomatic-speech-recognition100K<n<1M1 likes59 downloads23d agoHugging Face20bpop /spite-CV16-TP9B Spite Dataset Pseudolabeled speech translation data with quality annotations from multiple metrics. This version uses transcripts from Common Voice 16.1 and translations from Tower-Plus-9B. Configs en_de en_es en_fr en_it en_ko en_nl en_pt en_ru en_zh Usage from datasets import load_dataset ds = load_dataset("bpop/spite-CV16-Euro9B", "en_pt") tabulartranslation1M<n<10M0 likes46 downloads7mo agoHugging Face21Yehor /cv10-uk-testset-clean The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦 Overview This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios. All audios have been checked by a human to be sure that they are correct. This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk Community Discord: https://bit.ly/discord-uds Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.audioautomatic-speech-recognition1K<n<10K3 likes43 downloads2y agoHugging Face22bilguun /cv-mn-24.0 cv-mn-24.0 Mongolian subset of the Mozilla Common Voice speech recognition dataset. Dataset Statistics Total samples: 6,018Total duration: 9h 7m 2s (9.12 h) Per-split breakdown Split Samples Total Duration Avg Duration train 2,188 3h 7m 34s (3.13 h) 5.14 s validation 1,896 2h 54m 41s (2.91 h) 5.53 s test 1,934 3h 4m 47s (3.08 h) 5.73 s audioautomatic-speech-recognition1K<n<10K1 likes41 downloads5mo agoHugging Face23tiny-aya-translate /cv-tr-eval Common Voice Turkish Eval 4,825 Turkish test clips (~45 MB, 16 kHz) in the Mozilla Common Voice schema: transcription, duration, up_votes / down_votes, and the age / gender / accent speaker attributes. Schema in the YAML header above. An evaluation-only Turkish counterpart to lahaja-eval; never trained on. Used to sanity-check Turkish ASR quality on real human speech, which matters here because the v0.3 training corpus is entirely synthetic TTS and the released model is… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/cv-tr-eval.audioautomatic-speech-recognition1K<n<10K0 likes35 downloads2mo agoHugging Face24bpop /spite-CV16-Euro9B Spite Dataset Pseudolabeled speech translation data with quality annotations from multiple metrics. This version uses transcripts from Common Voice 16.1 and translations from EuroLLM-9B-Instruct. Configs en_de en_es en_fr en_it en_ko en_nl en_pt en_ru en_zh Usage from datasets import load_dataset ds = load_dataset("bpop/spite-CV16-Euro9B", "en_pt") tabulartranslation1M<n<10M0 likes32 downloads7mo agoHugging Face25Trelis /cv-en-scripted-test-500 Common Voice English Scripted Test Set — 500 clips n = 500 utterances · private eval set for ASR benchmarking Source Derived from Mozilla Common Voice Scripted Speech 25.0 — English (test split), downloaded via the Mozilla Data Collective API (dataset ID cmndapwry02jnmh07dyo46mot, 94 GB tarball). Construction Starting from the full CV 25.0 English test split (16,398 rows), a stratified 500-clip subset was produced using the same recipe as… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/cv-en-scripted-test-500.audioautomatic-speech-recognitionn<1K0 likes26 downloads4mo agoHugging Face26cvxhull /stt-calibration STT Calibration Dataset Tiny calibration dataset for PersonalAssistant STT service. Used on first run to auto-tune speculative pre-transcription parameters. Contents File Duration Size Purpose short.wav 3.5s 110KB RTF measurement + VAD onset latency long.wav 23.3s 729KB Split quality calibration (whole vs split comparison) very_long.wav 56.8s 1.8MB Multi-split calibration (find minimum safe split interval) manifest.json - 2KB Sample metadata + reference… See the full description on the dataset page: https://huggingface.co/datasets/cvxhull/stt-calibration.audioautomatic-speech-recognitionn<1K0 likes22 downloads7mo agoHugging Face27Veronica1NW /cv17_sw_kenyan_sample Common Voice 17.0 — Swahili (Kenyan Sample) This dataset is a filtered sample of the Mozilla Common Voice 17.0 corpus, focusing on Swahili (sw) speech with Kenyan voices. It has been subsetted for experimentation and prototyping in ASR (Automatic Speech Recognition) models targeting speech-impaired users in Kenya, covering Kenyan English and Kiswahili. Dataset Summary Language: Kiswahili (Swahili, sw) Accent/Region: Kenyan speakers Domain: Conversational… See the full description on the dataset page: https://huggingface.co/datasets/Veronica1NW/cv17_sw_kenyan_sample.audioautomatic-speech-recognition1K<n<10K0 likes21 downloads1y agoHugging Face28instinct-org /cv_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/cv_chunked Aligned dataset: instinct-org/cv_chunked_nfa_aligned Rows: 71097 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_nfa_aligned.textautomatic-speech-recognition10K<n<100K0 likes9 downloads1mo agoHugging Face29instinct-org /cv_chunkedgated cv_chunked This is a gated Uzbek chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked.audioautomatic-speech-recognition10K<n<100K0 likes7 downloads1mo agoHugging Face30instinct-org /cv_chunked_speech_restorisedgated cv_chunked_speech_restorised This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_speech_restorised.audioautomatic-speech-recognition10K<n<100K0 likes6 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.