CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saeedzou /common-voice-17-en-age-gender-accentaudio100K<n<1M0 likes688 downloads2mo agoHugging Face02saeedzou /persianvox_2_rawgatedaudio0 likes519 downloads1mo agoHugging Face03saeedzou /common-voice-17-en-age-genderaudio100K<n<1M0 likes496 downloads2mo agoHugging Face04saeedzou /persianvox_2_audiogatedaudio1K<n<10K0 likes487 downloads1mo agoHugging Face05saeedzou /vctk-48khzgated Dataset Card for VCTK (48kHz) This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset. A companion version resampled to 16kHz is also available: saeedzou/vctk-16khz. Dataset Summary This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker reads out… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-48khz.audioautomatic-speech-recognition10K<n<100K1 likes333 downloads2mo agoHugging Face06saeedzou /vctk-16khz Dataset Card for VCTK (16kHz) This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset. A companion version at the original 48kHz sample rate is also available: saeedzou/vctk-48khz. Dataset Summary This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-16khz.audioautomatic-speech-recognition10K<n<100K0 likes291 downloads2mo agoHugging Face07laion /majestrino-1.00-16xk5-sae-features Majestrino 1.00 SAE — Feature Audio Samples (16x, k=5) Top-2000 activating audio samples for each feature in the Majestrino 1.00 SAE. Overview Metric Value SAE Architecture 16x expansion, k=5, d_model=768 Total Features 12,288 Alive Features 10,684 Audio per Feature Up to 2,000 highest-activating Audio Format Opus (24 kbps OGG container) Total TAR Files 1069 Source Dataset laion/majestrino-data File Structure Each TAR file… See the full description on the dataset page: https://huggingface.co/datasets/laion/majestrino-1.00-16xk5-sae-features.audioaudio-classification10M<n<100M0 likes272 downloads6mo agoHugging Face08saeeew /JP-HomophoneBench JP-HomophoneBench A deterministic Japanese ASR benchmark index for separating eight error/disambiguation classes: exact_homophone near_homophone voicing long_vowel geminate moraic_nasal pitch_accent semantic_only Important design rule This repository is metadata-first. Source audio is not redistributed by default. Each row stores source repository/config/split/row identifiers so audio can be rehydrated under the original source license. exact_homophone and… See the full description on the dataset page: https://huggingface.co/datasets/saeeew/JP-HomophoneBench.audioautomatic-speech-recognitionn<1K0 likes191 downloads25d agoHugging Face09saeedzou /iemocap-original-wavlm-large-layer-9-temporalaudio1K<n<10K0 likes168 downloads2mo agoHugging Face10saeedzou /iemocap-vc-wavlm-large-layer-9-temporalaudio1K<n<10K0 likes165 downloads2mo agoHugging Face11saeedzou /e-daic-ai-controlledgatedaudio1K<n<10K0 likes162 downloads2mo agoHugging Face12saeedzou /iemocap-original-wavlm-layer-6-temporalaudio1K<n<10K0 likes138 downloads2mo agoHugging Face13saeedzou /iemocap-vc-wavlm-layer-6-temporalaudio1K<n<10K0 likes119 downloads2mo agoHugging Face14saeedzou /DAIC-WOZgatedaudio10K<n<100K0 likes113 downloads2mo agoHugging Face15saeedzou /common-voice-17-en-age-gender-sampledaudio10K<n<100K0 likes108 downloads2mo agoHugging Face16saeedzou /persianvox_allgated PersianVox: A Prosody-Aware Approach for Speech Dataset Generation from In-the-Wild Data PersianVox-All is the larger, less-filtered counterpart of PersianVox: a multi-speaker Persian (Farsi) speech corpus automatically mined from in-the-wild unlabeled data. It contains every utterance that passed language and speech-quality (MOS) filtering, without the additional dual-ASR transcript-agreement filtering applied to the main PersianVox release. It is therefore substantially… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/persianvox_all.audiotext-to-speech1M<n<10M0 likes100 downloads8d agoHugging Face17saeedzou /common-voice-17-en-age-gender-accent-sampledaudio10K<n<100K0 likes97 downloads2mo agoHugging Face18saeedzou /persianvoxgated PersianVox: A Prosody-Aware Approach for Speech Dataset Generation from In-the-Wild Data PersianVox is a 2,400-hour, multi-speaker Persian (Farsi) speech corpus automatically mined from in-the-wild unlabeled data. It is, to date, the largest open-source speech resource for Persian, built to support zero-shot text-to-speech (TTS) research and other speech tasks in low-resource-language settings. Dataset Summary Advancement of zero-shot text-to-speech… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/persianvox.audiotext-to-speech100K<n<1M0 likes74 downloads8d agoHugging Face19saeedzou /yodas2-en-000-116-00000000-emotion-filtered YODAS2 Emotion Dataset Pipeline This project builds an automatically labeled speech emotion dataset from the English portion of YODAS2. The pipeline combines speaker diarization, voice activity detection, speech segmentation, and predictions from five pretrained speech emotion recognition models. The resulting labels are filtered using model agreement and then downsampled to reduce the strong class imbalance in the source data. Source Data The source dataset is… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/yodas2-en-000-116-00000000-emotion-filtered.audio10K<n<100K0 likes73 downloads2mo agoHugging Face20saeedzou /sep28k-fluencybank-stutter-datasetaudio10K<n<100K0 likes69 downloads2mo agoHugging Face21saeedzou /sep28k-wavlm-layer-9audio10K<n<100K0 likes66 downloads2mo agoHugging Face22saeedzou /persianvox_rawgatedaudio10K<n<100K0 likes61 downloads16d agoHugging Face23saeedzou /DataSEDaudion<1K0 likes61 downloads10d agoHugging Face24saeedzou /saarbruecken-voice-database-16khzgatedaudio10K<n<100K0 likes59 downloads2mo agoHugging Face25saefewfe /asfgtrfyjyuaudion<1K0 likes46 downloads22d agoHugging Face26saeedzou /iemocap-vc_vctk_2-wavlm-large-layer-9-mosaudio10K<n<100K0 likes43 downloads2mo agoHugging Face27saeedzou /iemocap-vc_vctkaudio1K<n<10K0 likes40 downloads2mo agoHugging Face28saeedzou /iemocap-vc_vctk_2audio10K<n<100K0 likes40 downloads2mo agoHugging Face29saeedzou /iemocap-vc_vctk_2-wavlm-large-layer-9audio10K<n<100K0 likes35 downloads2mo agoHugging Face30saeedzou /English_Accent_DataSet_Filtered_400audio1K<n<10K0 likes35 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.