datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.CoVoGER
CoVoGER: A Multilingual Multitask Benchmark for Speech-to-text Generative Error Correction with Large Language Models
Dataset Description
Large language models (LLMs) can rewrite the N-best hypotheses from a speech-to-text model, often fixing recognition or translation errors that traditional rescoring cannot. Yet research on generative error correction (GER) has been focusing on monolingual automatic speech recognition (ASR), leaving its multilingual and multitask… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/CoVoGER.tajik-asr-corpus-v3
tajik-asr-corpus-v3
1,071 hours of Tajik ASR training data: 41 Tajik YouTube channels (~1,059 h, machine-labeled)
plus FLEURS tg_tj (11.8 h, gold). This is the dataset behind
Peacockery/omni-ctc-300m-tajik
(16.9% WER on FLEURS test, 37.6% on held-out conversational speech).
Layout
Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=tgk_Cyrl/.
Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list),
and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v3.tajik-asr-corpus-v0
Tajik ASR Corpus v0
Deduplicated Tajik automatic speech recognition corpus assembled from FLEURS-derived
speech data, Mozilla Common Voice 25 Tajik, and Muhtasham Tajik ASR augmented data.
Format
Each split has a data.tsv and an audio/ directory.
TSV columns:
id
audio_filename
raw_transcription
transcription
characters
audio_bytes
source
source_id
duplicate_count
tajik_asr_combined.sqlite mirrors the TSV rows and includes normalized_text,
source_split, and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v0.tajik-asr-youtube
tajik-asr-youtube
Every Tajik-language clip from this project's YouTube scrape — 41 channels of news, talk
shows, podcasts, audiobooks, and learning content — with machine transcripts and the
verification scores left in as columns instead of applied as a filter. Pick your own
quality threshold; the training corpus this project actually ships
(tajik-asr-corpus-v3)
is the gated subset.
Layout
Parquet shards under data/ with an audio struct column (16 kHz mono FLAC… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-youtube.farsi-asr-corpus-v4
farsi-asr-corpus-v4
985 hours of Farsi ASR training data across seven corpora: Common Voice 25 (324 h), Thomcles Farsi speech (298 h), Farsi YouTube (205 h), Mana TTS (97 h), Neyshekar (37 h), FLEURS fa_ir (13 h), WorldSpeech (11 h). This is the dataset behind Peacockery/omni-ctc-300m-farsi (8.5% WER on FLEURS test).
Layout
Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=fas_Arab/. Each row holds text (the normalized label)… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/farsi-asr-corpus-v4.mozilla-common-voice-spontaneous-speech-asr-shared-task
Mozilla Common Voice Spontaneous Speech ASR Shared Task
This repository combines the Mozilla Data Collective Common Voice spontaneous speech ASR shared-task
train/dev and test archives in one place.
Locales present across the combined train/dev and test packages: ady, aln, bas, bew, bxk,
cgg, el-CY, hch, kbd, kcn, koo, led, lke, lth, meh, mmc, pne, qxp, ruc,
rwm, sco, tob, top, ttj, ukv, ush.
Split package
Mozilla Data Collective dataset ID
Hub archive
Original MDC archive… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/mozilla-common-voice-spontaneous-speech-asr-shared-task.peaky-blinders-learning-purpose-only
TTS Dataset - Peaky Blinders
This dataset contains audio segments with transcriptions from Peaky Blinders for Text-to-Speech training.
Dataset Generation
This dataset was generated using the TTS-Dataset-Maker pipeline, which provides:
Silero VAD-based silence removal - Removes long silences while preserving natural speech gaps
DeepFilterNet denoising - CPU-optimized audio denoising with gentle attenuation (15dB)
AssemblyAI transcription - High-quality speech-to-text with… See the full description on the dataset page: https://huggingface.co/datasets/ahk-d/peaky-blinders-learning-purpose-only.georgian-asr-corpus-v0
georgian-asr-corpus-v0
145.3 hours of Georgian ASR training data: 92,185 clips across FLEURS ka_ge and Common Voice Georgian (scripted 25.0 and spontaneous 3.0, via the Mozilla Data Collective). Splits: train 64,633 / dev 13,456 / test 14,096.
Layout
Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=kat_Geor/. Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list), and audio_size (sample count).… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/georgian-asr-corpus-v0.speechneyshekar-v3-asr-aligned
Neyshekar v3 ASR-Aligned
This is a repaired subset of Neyshekar v3 for Persian ASR work. The public v3
archive contains real audio and real transcripts, but the downloaded
dataset.json filename-to-text mapping does not align for the checked samples.
This export keeps only audio clips whose transcript could be recovered by
matching multiple ASR hypotheses back to the original Neyshekar transcript pool.
It is useful as a curated ASR training/evaluation candidate set, with the… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/neyshekar-v3-asr-aligned.farsi-asr-wer35-fastconformer
Farsi ASR WER35 FastConformer
This dataset contains Farsi/Farsi ASR utterances curated with NVIDIA NeMo Curator using nvidia/stt_fa_fastconformer_hybrid_large.
The uploaded training data is stored as WebDataset TAR shards because Hugging Face recommends WebDataset archives for large-scale audio datasets. The local artifact also includes a NeMo ASR manifest at manifests/train_manifest.jsonl.
Curation
Language: Farsi/Farsi (fa)
Audio: FLAC, 16 kHz mono
Source run:… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/farsi-asr-wer35-fastconformer.
