datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common-voice-17-en-age-gender-accentpersianvox_2_rawcommon-voice-17-en-age-genderpersianvox_2_audiovctk-48khz
Dataset Card for VCTK (48kHz)
This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset.
A companion version resampled to 16kHz is also available: saeedzou/vctk-16khz.
Dataset Summary
This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker reads out… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-48khz.vctk-16khz
Dataset Card for VCTK (16kHz)
This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset.
A companion version at the original 48kHz sample rate is also available: saeedzou/vctk-48khz.
Dataset Summary
This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-16khz.majestrino-1.00-16xk5-sae-features
Majestrino 1.00 SAE — Feature Audio Samples (16x, k=5)
Top-2000 activating audio samples for each feature in the
Majestrino 1.00 SAE.
Overview
Metric
Value
SAE Architecture
16x expansion, k=5, d_model=768
Total Features
12,288
Alive Features
10,684
Audio per Feature
Up to 2,000 highest-activating
Audio Format
Opus (24 kbps OGG container)
Total TAR Files
1069
Source Dataset
laion/majestrino-data
File Structure
Each TAR file… See the full description on the dataset page: https://huggingface.co/datasets/laion/majestrino-1.00-16xk5-sae-features.JP-HomophoneBench
JP-HomophoneBench
A deterministic Japanese ASR benchmark index for separating eight error/disambiguation classes:
exact_homophone
near_homophone
voicing
long_vowel
geminate
moraic_nasal
pitch_accent
semantic_only
Important design rule
This repository is metadata-first. Source audio is not redistributed by default. Each row stores source repository/config/split/row identifiers so audio can be rehydrated under the original source license.
exact_homophone and… See the full description on the dataset page: https://huggingface.co/datasets/saeeew/JP-HomophoneBench.iemocap-original-wavlm-large-layer-9-temporaliemocap-vc-wavlm-large-layer-9-temporale-daic-ai-controllediemocap-original-wavlm-layer-6-temporaliemocap-vc-wavlm-layer-6-temporalDAIC-WOZcommon-voice-17-en-age-gender-sampledpersianvox_all
PersianVox: A Prosody-Aware Approach for Speech Dataset Generation from In-the-Wild Data
PersianVox-All is the larger, less-filtered counterpart of PersianVox: a multi-speaker Persian (Farsi) speech corpus automatically mined from in-the-wild unlabeled data. It contains every utterance that passed language and speech-quality (MOS) filtering, without the additional dual-ASR transcript-agreement filtering applied to the main PersianVox release. It is therefore substantially… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/persianvox_all.common-voice-17-en-age-gender-accent-sampledpersianvox
PersianVox: A Prosody-Aware Approach for Speech Dataset Generation from In-the-Wild Data
PersianVox is a 2,400-hour, multi-speaker Persian (Farsi) speech corpus automatically mined from in-the-wild unlabeled data. It is, to date, the largest open-source speech resource for Persian, built to support zero-shot text-to-speech (TTS) research and other speech tasks in low-resource-language settings.
Dataset Summary
Advancement of zero-shot text-to-speech… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/persianvox.yodas2-en-000-116-00000000-emotion-filtered
YODAS2 Emotion Dataset Pipeline
This project builds an automatically labeled speech emotion dataset from the English portion of YODAS2.
The pipeline combines speaker diarization, voice activity detection, speech segmentation, and predictions from five pretrained speech emotion recognition models. The resulting labels are filtered using model agreement and then downsampled to reduce the strong class imbalance in the source data.
Source Data
The source dataset is… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/yodas2-en-000-116-00000000-emotion-filtered.sep28k-fluencybank-stutter-datasetsep28k-wavlm-layer-9persianvox_rawDataSEDsaarbruecken-voice-database-16khzasfgtrfyjyuiemocap-vc_vctk_2-wavlm-large-layer-9-mosiemocap-vc_vctkiemocap-vc_vctk_2iemocap-vc_vctk_2-wavlm-large-layer-9English_Accent_DataSet_Filtered_400
