CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Peacockery /common-voice-scripted-speech-26 Common Voice Scripted Speech A row-normalized multilingual ASR dataset built from Mozilla Data Collective Common Voice Scripted Speech. Each upstream archive is converted to appendable parquet shards under data/<upstream_split>/, one shard per source archive and split, with audio bytes embedded in an audio struct column. Status Manifest languages: 60 Languages uploaded: 18 Columns audio (bytes, path) sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.tabularautomatic-speech-recognition100K<n<1M0 likes2k downloads3mo agoHugging Face02PeacefulData /CoVoGER CoVoGER: A Multilingual Multitask Benchmark for Speech-to-text Generative Error Correction with Large Language Models Dataset Description Large language models (LLMs) can rewrite the N-best hypotheses from a speech-to-text model, often fixing recognition or translation errors that traditional rescoring cannot. Yet research on generative error correction (GER) has been focusing on monolingual automatic speech recognition (ASR), leaving its multilingual and multitask… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/CoVoGER.automatic-speech-recognition1 likes1.4k downloads6mo agoHugging Face03Peacockery /tajik-asr-corpus-v3 tajik-asr-corpus-v3 1,071 hours of Tajik ASR training data: 41 Tajik YouTube channels (~1,059 h, machine-labeled) plus FLEURS tg_tj (11.8 h, gold). This is the dataset behind Peacockery/omni-ctc-300m-tajik (16.9% WER on FLEURS test, 37.6% on held-out conversational speech). Layout Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=tgk_Cyrl/. Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list), and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v3.textautomatic-speech-recognition100K<n<1M3 likes131 downloads4mo agoHugging Face04Peacockery /tajik-asr-corpus-v0 Tajik ASR Corpus v0 Deduplicated Tajik automatic speech recognition corpus assembled from FLEURS-derived speech data, Mozilla Common Voice 25 Tajik, and Muhtasham Tajik ASR augmented data. Format Each split has a data.tsv and an audio/ directory. TSV columns: id audio_filename raw_transcription transcription characters audio_bytes source source_id duplicate_count tajik_asr_combined.sqlite mirrors the TSV rows and includes normalized_text, source_split, and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v0.audioautomatic-speech-recognition1K<n<10K1 likes104 downloads4mo agoHugging Face05Peacockery /tajik-asr-youtube tajik-asr-youtube Every Tajik-language clip from this project's YouTube scrape — 41 channels of news, talk shows, podcasts, audiobooks, and learning content — with machine transcripts and the verification scores left in as columns instead of applied as a filter. Pick your own quality threshold; the training corpus this project actually ships (tajik-asr-corpus-v3) is the gated subset. Layout Parquet shards under data/ with an audio struct column (16 kHz mono FLAC… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-youtube.tabularautomatic-speech-recognition100K<n<1M0 likes56 downloads4mo agoHugging Face06Peacockery /farsi-asr-corpus-v4 farsi-asr-corpus-v4 985 hours of Farsi ASR training data across seven corpora: Common Voice 25 (324 h), Thomcles Farsi speech (298 h), Farsi YouTube (205 h), Mana TTS (97 h), Neyshekar (37 h), FLEURS fa_ir (13 h), WorldSpeech (11 h). This is the dataset behind Peacockery/omni-ctc-300m-farsi (8.5% WER on FLEURS test). Layout Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=fas_Arab/. Each row holds text (the normalized label)… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/farsi-asr-corpus-v4.textautomatic-speech-recognition100K<n<1M0 likes49 downloads4mo agoHugging Face07Peacockery /mozilla-common-voice-spontaneous-speech-asr-shared-task Mozilla Common Voice Spontaneous Speech ASR Shared Task This repository combines the Mozilla Data Collective Common Voice spontaneous speech ASR shared-task train/dev and test archives in one place. Locales present across the combined train/dev and test packages: ady, aln, bas, bew, bxk, cgg, el-CY, hch, kbd, kcn, koo, led, lke, lth, meh, mmc, pne, qxp, ruc, rwm, sco, tob, top, ttj, ukv, ush. Split package Mozilla Data Collective dataset ID Hub archive Original MDC archive… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/mozilla-common-voice-spontaneous-speech-asr-shared-task.textautomatic-speech-recognition10K<n<100K0 likes43 downloads4mo agoHugging Face08ahk-d /peaky-blinders-learning-purpose-only TTS Dataset - Peaky Blinders This dataset contains audio segments with transcriptions from Peaky Blinders for Text-to-Speech training. Dataset Generation This dataset was generated using the TTS-Dataset-Maker pipeline, which provides: Silero VAD-based silence removal - Removes long silences while preserving natural speech gaps DeepFilterNet denoising - CPU-optimized audio denoising with gentle attenuation (15dB) AssemblyAI transcription - High-quality speech-to-text with… See the full description on the dataset page: https://huggingface.co/datasets/ahk-d/peaky-blinders-learning-purpose-only.audiotext-to-speechn<1K0 likes36 downloads1y agoHugging Face09Peacockery /georgian-asr-corpus-v0 georgian-asr-corpus-v0 145.3 hours of Georgian ASR training data: 92,185 clips across FLEURS ka_ge and Common Voice Georgian (scripted 25.0 and spontaneous 3.0, via the Mozilla Data Collective). Splits: train 64,633 / dev 13,456 / test 14,096. Layout Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=kat_Geor/. Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list), and audio_size (sample count).… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/georgian-asr-corpus-v0.textautomatic-speech-recognition10K<n<100K0 likes24 downloads4mo agoHugging Face10peanut999 /speechtextautomatic-speech-recognitionn<1K0 likes8 downloads2y agoHugging Face11Peacockery /neyshekar-v3-asr-aligned Neyshekar v3 ASR-Aligned This is a repaired subset of Neyshekar v3 for Persian ASR work. The public v3 archive contains real audio and real transcripts, but the downloaded dataset.json filename-to-text mapping does not align for the checked samples. This export keeps only audio clips whose transcript could be recovered by matching multiple ASR hypotheses back to the original Neyshekar transcript pool. It is useful as a curated ASR training/evaluation candidate set, with the… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/neyshekar-v3-asr-aligned.audioautomatic-speech-recognition2 likes7 downloads5mo agoHugging Face12Peacockery /farsi-asr-wer35-fastconformer Farsi ASR WER35 FastConformer This dataset contains Farsi/Farsi ASR utterances curated with NVIDIA NeMo Curator using nvidia/stt_fa_fastconformer_hybrid_large. The uploaded training data is stored as WebDataset TAR shards because Hugging Face recommends WebDataset archives for large-scale audio datasets. The local artifact also includes a NeMo ASR manifest at manifests/train_manifest.jsonl. Curation Language: Farsi/Farsi (fa) Audio: FLAC, 16 kHz mono Source run:… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/farsi-asr-wer35-fastconformer.automatic-speech-recognition0 likes5 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.