CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01resproj007 /pathological_speech Pathological Speech (TORGO + UA-Speech + LibriSpeech Normal) Mixed-corpus speech dataset for training and evaluating controllable speech-synthesis and severity-classification models. Three corpora are merged with unified metadata so a single model can learn severity- and gender-conditioned generation without confounds. Splits (speaker-disjoint since 2026-09-14) Split Rows Bytes (parquet) What it is train 37704 5,500,328,387 every clip of every speaker… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/pathological_speech.audio10K<n<100K0 likes1.2k downloads8d agoHugging Face02K-University-AIED /Pathological-child-voice Speech Dataset for AI-Based Language Assessment in Children The "Speech Database of Typically Developing and Speech-Impaired Children" is an open speech dataset designed to support the development of AI-based language assessment systems. It contains speech samples from children aged 2 to 9 who are either typically developing or have reduced consonant articulation accuracy. This dataset is based on standardized Korean articulation tools: APAC (Articulation and Phonology… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/Pathological-child-voice.audion<1K0 likes57 downloads27d agoHugging Face03The-Nature-of-Reality /THE-PATH-TO-THE-NEW-WORLDaudion<1K1 likes31 downloads1y agoHugging Face04K-Univ /Pathological-child-voice Speech Dataset for AI-Based Language Assessment in Children The "Speech Database of Typically Developing and Speech-Impaired Children" is an open speech dataset designed to support the development of AI-based language assessment systems. It contains speech samples from children aged 2 to 9 who are either typically developing or have reduced consonant articulation accuracy. This dataset is based on standardized Korean articulation tools: APAC (Articulation and Phonology Assessment… See the full description on the dataset page: https://huggingface.co/datasets/K-Univ/Pathological-child-voice.audion<1K0 likes30 downloads1y agoHugging Face05resproj007 /original_pathological_speechaudion<1K0 likes24 downloads7mo agoHugging Face06resproj007 /orpheus_tts_knn_vc_pathological Combined Orpheus TTS 3B + Same-Speaker KNN Voice Conversion Dataset Dataset Overview This dataset contains fine-tuned Orpheus TTS 3B synthetic speech enhanced with same-speaker KNN voice conversion a Enhancement Method: Orpheus TTS 3B LoRA fine-tuned synthetic speech → Same-speaker KNN voice conversion using real audio references from the same speaker Dataset Statistics Total Samples: 759 Total Duration: 2583.38 seconds (43.06 minutes) Speakers: 8 speakers… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/orpheus_tts_knn_vc_pathological.audion<1K0 likes19 downloads1y agoHugging Face07resproj007 /spark_pathological_synthetic_data Combined Fine-tuned Spark TTS Synthetic Speech Dataset Dataset Statistics Total Samples: 785 Total Duration: 0.8 hours Speakers: 8 Corpora: TORGO, UA-Speech, LibriSpeech Audio Format: 16kHz WAV Data Format Each sample contains: audio: Audio array with sampling_rate (16kHz) text: Original transcript text speaker_id: Speaker identifier (F04, M02, FC02, MC01, F02, M04, 211, 4014) corpus: Source corpus (TORGO, UA-Speech, LibriSpeech) condition: Speaker condition… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/spark_pathological_synthetic_data.audion<1K0 likes15 downloads1y agoHugging Face08saeedzou /femh-pathological-voicegatedaudio1K<n<10K0 likes15 downloads2mo agoHugging Face09resproj007 /spark_tts_knn_vc_pathological Dataset Overview Total Samples: 785 Total Duration: 3006.46 seconds (50.11 minutes) Speakers: 8 speakers Corpora: TORGO, UA-Speech, LibriSpeech Sample Rate: 16kHz (KNN-VC output rate, native Spark compatibility) Audio Format: WAV audion<1K0 likes10 downloads1y agoHugging Face10WiktorJakubowski /MELD-videos-absolute-pathsaudio10K<n<100K0 likes9 downloads1y agoHugging Face11emilykang /pathology_trainaudio1K<n<10K2 likes8 downloads3y agoHugging Face12emilykang /pathology_testaudion<1K1 likes8 downloads3y agoHugging Face13resproj007 /sesame_pathological_synthetic_data Combined Fine-tuned Sesame CSM 1B Synthetic Speech Dataset Data Format Each sample contains: audio: Audio array with sampling_rate (24kHz) text: Original transcript text speaker_id: Speaker identifier (F04, M02, FC02, MC01, F02, M04, 211, 4014) corpus: Source corpus (TORGO, UA-Speech, LibriSpeech) condition: Speaker condition (Dysarthric, Healthy) model_name: Fine-tuned model name model_type: "sesame_csm_1b_adapter" base_model: Base model used (unsloth/csm-1b)… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/sesame_pathological_synthetic_data.audion<1K0 likes7 downloads1y agoHugging Face14resproj007 /sesame_tts_knn_vc_pathological Dataset Overview Total Samples: 759 Total Duration: 2892.82 seconds (48.21 minutes) Speakers: 8 speakers Corpora: TORGO, UA-Speech, LibriSpeech Sample Rate: 16kHz (KNN-VC output rate) Audio Format: WAV audion<1K0 likes7 downloads1y agoHugging Face15resproj007 /pathological_knn_vc Combined Same-Speaker KNN Voice Conversion Dataset Dataset Statistics Total Samples: 799 Total Duration: 0.4 hours Speakers: 8 Corpora: TORGO, UA-Speech, LibriSpeech Audio Format: 16kHz WAV Baseline TTS: resproj007/baseline_orpheus_3b Speaker Breakdown Speaker Name Corpus Condition Gender Samples Duration FC02 TORGO Healthy Female TORGO Healthy Female 100 146.0s M04 UA-Speech Male UA-Speech Dysarthric Male 200 245.7s M02 TORGO Dysarthric Male… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/pathological_knn_vc.audion<1K0 likes5 downloads1y agoHugging Face16resproj007 /orpheus_pathological_synthetic_data Combined Fine-tuned Orpheus TTS 3B Synthetic Speech Dataset Data Format Each sample contains: audio: Audio array with sampling_rate (24kHz) text: Original transcript text speaker_id: Speaker identifier (F04, M02, FC02, MC01, F02, M04, 211, 4014) corpus: Source corpus (TORGO, UA-Speech, LibriSpeech) condition: Speaker condition (Dysarthric, Healthy) model_name: Fine-tuned model name model_type: "orpheus_3b_adapter" base_model: Base model used (unsloth/orpheus-3b-0.1-ft)… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/orpheus_pathological_synthetic_data.audion<1K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.