CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saeedzou /vctk-48khzgated Dataset Card for VCTK (48kHz) This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset. A companion version resampled to 16kHz is also available: saeedzou/vctk-16khz. Dataset Summary This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker reads out… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-48khz.audioautomatic-speech-recognition10K<n<100K1 likes333 downloads2mo agoHugging Face02saeedzou /vctk-16khz Dataset Card for VCTK (16kHz) This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset. A companion version at the original 48kHz sample rate is also available: saeedzou/vctk-48khz. Dataset Summary This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-16khz.audioautomatic-speech-recognition10K<n<100K0 likes291 downloads2mo agoHugging Face03saeeew /JP-HomophoneBench JP-HomophoneBench A deterministic Japanese ASR benchmark index for separating eight error/disambiguation classes: exact_homophone near_homophone voicing long_vowel geminate moraic_nasal pitch_accent semantic_only Important design rule This repository is metadata-first. Source audio is not redistributed by default. Each row stores source repository/config/split/row identifiers so audio can be rehydrated under the original source license. exact_homophone and… See the full description on the dataset page: https://huggingface.co/datasets/saeeew/JP-HomophoneBench.audioautomatic-speech-recognitionn<1K0 likes191 downloads25d agoHugging Face04saeedzou /persianvox_allgated PersianVox: A Prosody-Aware Approach for Speech Dataset Generation from In-the-Wild Data PersianVox-All is the larger, less-filtered counterpart of PersianVox: a multi-speaker Persian (Farsi) speech corpus automatically mined from in-the-wild unlabeled data. It contains every utterance that passed language and speech-quality (MOS) filtering, without the additional dual-ASR transcript-agreement filtering applied to the main PersianVox release. It is therefore substantially… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/persianvox_all.audiotext-to-speech1M<n<10M0 likes100 downloads8d agoHugging Face05saeedzou /persianvoxgated PersianVox: A Prosody-Aware Approach for Speech Dataset Generation from In-the-Wild Data PersianVox is a 2,400-hour, multi-speaker Persian (Farsi) speech corpus automatically mined from in-the-wild unlabeled data. It is, to date, the largest open-source speech resource for Persian, built to support zero-shot text-to-speech (TTS) research and other speech tasks in low-resource-language settings. Dataset Summary Advancement of zero-shot text-to-speech… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/persianvox.audiotext-to-speech100K<n<1M0 likes74 downloads8d agoHugging Face06saeedzou /PersianVox_HBgated PersianVox_HB PersianVox_HB is a high-quality, multispeaker Persian speech dataset derived from audio recordings of the Holy Bible. The dataset is sourced from bible.com (PCB=49.85 hours, TPV=71.4 hours), bible.is (NMV=24.17 hours), and wordproject.org (PHB=81.92 hours). 📚 Dataset Summary Language: Persian (Farsi) Speakers: Multiple speakers Total Duration: 227.34 hours Recording Sources: bible.com (PCB: 49.85 hours, TPV: 71.4 hours) bible.is (NMV: 24.17 hours)… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/PersianVox_HB.audioautomatic-speech-recognition10K<n<100K6 likes7 downloads1y agoHugging Face07saeedzou /PersianVox_NMgated PersianVox_NM PersianVox_NM is a high-quality Persian speech dataset derived from article readings by a single female speaker. This subset is sourced from Nasle Mana Magazine, a publication dedicated to producing accessible content for the visually impaired. 📚 Dataset Summary Language: Persian (Farsi) Speaker: Single female voice Total Duration: 94.55 hours Recording Source: Articles from Nasle Mana Domain: Literary and informational prose Alignment Checked With:… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/PersianVox_NM.audioautomatic-speech-recognition10K<n<100K6 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.