datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vctk-48khz
Dataset Card for VCTK (48kHz)
This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset.
A companion version resampled to 16kHz is also available: saeedzou/vctk-16khz.
Dataset Summary
This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker reads out… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-48khz.vctk-16khz
Dataset Card for VCTK (16kHz)
This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset.
A companion version at the original 48kHz sample rate is also available: saeedzou/vctk-48khz.
Dataset Summary
This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-16khz.JP-HomophoneBench
JP-HomophoneBench
A deterministic Japanese ASR benchmark index for separating eight error/disambiguation classes:
exact_homophone
near_homophone
voicing
long_vowel
geminate
moraic_nasal
pitch_accent
semantic_only
Important design rule
This repository is metadata-first. Source audio is not redistributed by default. Each row stores source repository/config/split/row identifiers so audio can be rehydrated under the original source license.
exact_homophone and… See the full description on the dataset page: https://huggingface.co/datasets/saeeew/JP-HomophoneBench.persianvox_all
PersianVox: A Prosody-Aware Approach for Speech Dataset Generation from In-the-Wild Data
PersianVox-All is the larger, less-filtered counterpart of PersianVox: a multi-speaker Persian (Farsi) speech corpus automatically mined from in-the-wild unlabeled data. It contains every utterance that passed language and speech-quality (MOS) filtering, without the additional dual-ASR transcript-agreement filtering applied to the main PersianVox release. It is therefore substantially… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/persianvox_all.persianvox
PersianVox: A Prosody-Aware Approach for Speech Dataset Generation from In-the-Wild Data
PersianVox is a 2,400-hour, multi-speaker Persian (Farsi) speech corpus automatically mined from in-the-wild unlabeled data. It is, to date, the largest open-source speech resource for Persian, built to support zero-shot text-to-speech (TTS) research and other speech tasks in low-resource-language settings.
Dataset Summary
Advancement of zero-shot text-to-speech… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/persianvox.PersianVox_HB
PersianVox_HB
PersianVox_HB is a high-quality, multispeaker Persian speech dataset derived from audio recordings of the Holy Bible. The dataset is sourced from bible.com (PCB=49.85 hours, TPV=71.4 hours), bible.is (NMV=24.17 hours), and wordproject.org (PHB=81.92 hours).
📚 Dataset Summary
Language: Persian (Farsi)
Speakers: Multiple speakers
Total Duration: 227.34 hours
Recording Sources:
bible.com (PCB: 49.85 hours, TPV: 71.4 hours)
bible.is (NMV: 24.17 hours)… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/PersianVox_HB.PersianVox_NM
PersianVox_NM
PersianVox_NM is a high-quality Persian speech dataset derived from article readings by a single female speaker. This subset is sourced from Nasle Mana Magazine, a publication dedicated to producing accessible content for the visually impaired.
📚 Dataset Summary
Language: Persian (Farsi)
Speaker: Single female voice
Total Duration: 94.55 hours
Recording Source: Articles from Nasle Mana
Domain: Literary and informational prose
Alignment Checked With:… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/PersianVox_NM.
