CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Peacockery /common-voice-scripted-speech-26 Common Voice Scripted Speech A row-normalized multilingual ASR dataset built from Mozilla Data Collective Common Voice Scripted Speech. Each upstream archive is converted to appendable parquet shards under data/<upstream_split>/, one shard per source archive and split, with audio bytes embedded in an audio struct column. Status Manifest languages: 60 Languages uploaded: 18 Columns audio (bytes, path) sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.tabularautomatic-speech-recognition100K<n<1M0 likes2k downloads3mo agoHugging Face02Peacockery /tajik-asr-corpus-v3 tajik-asr-corpus-v3 1,071 hours of Tajik ASR training data: 41 Tajik YouTube channels (~1,059 h, machine-labeled) plus FLEURS tg_tj (11.8 h, gold). This is the dataset behind Peacockery/omni-ctc-300m-tajik (16.9% WER on FLEURS test, 37.6% on held-out conversational speech). Layout Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=tgk_Cyrl/. Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list), and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v3.textautomatic-speech-recognition100K<n<1M3 likes131 downloads4mo agoHugging Face03Peacockery /tajik-asr-youtube tajik-asr-youtube Every Tajik-language clip from this project's YouTube scrape — 41 channels of news, talk shows, podcasts, audiobooks, and learning content — with machine transcripts and the verification scores left in as columns instead of applied as a filter. Pick your own quality threshold; the training corpus this project actually ships (tajik-asr-corpus-v3) is the gated subset. Layout Parquet shards under data/ with an audio struct column (16 kHz mono FLAC… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-youtube.tabularautomatic-speech-recognition100K<n<1M0 likes56 downloads4mo agoHugging Face04Peacockery /farsi-asr-corpus-v4 farsi-asr-corpus-v4 985 hours of Farsi ASR training data across seven corpora: Common Voice 25 (324 h), Thomcles Farsi speech (298 h), Farsi YouTube (205 h), Mana TTS (97 h), Neyshekar (37 h), FLEURS fa_ir (13 h), WorldSpeech (11 h). This is the dataset behind Peacockery/omni-ctc-300m-farsi (8.5% WER on FLEURS test). Layout Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=fas_Arab/. Each row holds text (the normalized label)… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/farsi-asr-corpus-v4.textautomatic-speech-recognition100K<n<1M0 likes49 downloads4mo agoHugging Face05Peacockery /mozilla-common-voice-spontaneous-speech-asr-shared-task Mozilla Common Voice Spontaneous Speech ASR Shared Task This repository combines the Mozilla Data Collective Common Voice spontaneous speech ASR shared-task train/dev and test archives in one place. Locales present across the combined train/dev and test packages: ady, aln, bas, bew, bxk, cgg, el-CY, hch, kbd, kcn, koo, led, lke, lth, meh, mmc, pne, qxp, ruc, rwm, sco, tob, top, ttj, ukv, ush. Split package Mozilla Data Collective dataset ID Hub archive Original MDC archive… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/mozilla-common-voice-spontaneous-speech-asr-shared-task.textautomatic-speech-recognition10K<n<100K0 likes43 downloads4mo agoHugging Face06Peacockery /georgian-asr-corpus-v0 georgian-asr-corpus-v0 145.3 hours of Georgian ASR training data: 92,185 clips across FLEURS ka_ge and Common Voice Georgian (scripted 25.0 and spontaneous 3.0, via the Mozilla Data Collective). Splits: train 64,633 / dev 13,456 / test 14,096. Layout Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=kat_Geor/. Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list), and audio_size (sample count).… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/georgian-asr-corpus-v0.textautomatic-speech-recognition10K<n<100K0 likes24 downloads4mo agoHugging Face07peanut999 /speechtextautomatic-speech-recognitionn<1K0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.