datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.tajik-asr-corpus-v3
tajik-asr-corpus-v3
1,071 hours of Tajik ASR training data: 41 Tajik YouTube channels (~1,059 h, machine-labeled)
plus FLEURS tg_tj (11.8 h, gold). This is the dataset behind
Peacockery/omni-ctc-300m-tajik
(16.9% WER on FLEURS test, 37.6% on held-out conversational speech).
Layout
Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=tgk_Cyrl/.
Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list),
and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v3.tajik-asr-youtube
tajik-asr-youtube
Every Tajik-language clip from this project's YouTube scrape — 41 channels of news, talk
shows, podcasts, audiobooks, and learning content — with machine transcripts and the
verification scores left in as columns instead of applied as a filter. Pick your own
quality threshold; the training corpus this project actually ships
(tajik-asr-corpus-v3)
is the gated subset.
Layout
Parquet shards under data/ with an audio struct column (16 kHz mono FLAC… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-youtube.farsi-asr-corpus-v4
farsi-asr-corpus-v4
985 hours of Farsi ASR training data across seven corpora: Common Voice 25 (324 h), Thomcles Farsi speech (298 h), Farsi YouTube (205 h), Mana TTS (97 h), Neyshekar (37 h), FLEURS fa_ir (13 h), WorldSpeech (11 h). This is the dataset behind Peacockery/omni-ctc-300m-farsi (8.5% WER on FLEURS test).
Layout
Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=fas_Arab/. Each row holds text (the normalized label)… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/farsi-asr-corpus-v4.mozilla-common-voice-spontaneous-speech-asr-shared-task
Mozilla Common Voice Spontaneous Speech ASR Shared Task
This repository combines the Mozilla Data Collective Common Voice spontaneous speech ASR shared-task
train/dev and test archives in one place.
Locales present across the combined train/dev and test packages: ady, aln, bas, bew, bxk,
cgg, el-CY, hch, kbd, kcn, koo, led, lke, lth, meh, mmc, pne, qxp, ruc,
rwm, sco, tob, top, ttj, ukv, ush.
Split package
Mozilla Data Collective dataset ID
Hub archive
Original MDC archive… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/mozilla-common-voice-spontaneous-speech-asr-shared-task.georgian-asr-corpus-v0
georgian-asr-corpus-v0
145.3 hours of Georgian ASR training data: 92,185 clips across FLEURS ka_ge and Common Voice Georgian (scripted 25.0 and spontaneous 3.0, via the Mozilla Data Collective). Splits: train 64,633 / dev 13,456 / test 14,096.
Layout
Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=kat_Geor/. Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list), and audio_size (sample count).… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/georgian-asr-corpus-v0.speech
