datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pathological_speech
Pathological Speech (TORGO + UA-Speech + LibriSpeech Normal)
Mixed-corpus speech dataset for training and evaluating controllable
speech-synthesis and severity-classification models. Three corpora are merged
with unified metadata so a single model can learn severity- and
gender-conditioned generation without confounds.
Splits (speaker-disjoint since 2026-09-14)
Split
Rows
Bytes (parquet)
What it is
train
37704
5,500,328,387
every clip of every speaker… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/pathological_speech.Pathological-child-voice
Speech Dataset for AI-Based Language Assessment in Children
The "Speech Database of Typically Developing and Speech-Impaired Children" is an open speech dataset designed to support the development of AI-based language assessment systems. It contains speech samples from children aged 2 to 9 who are either typically developing or have reduced consonant articulation accuracy.
This dataset is based on standardized Korean articulation tools:
APAC (Articulation and Phonology… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/Pathological-child-voice.THE-PATH-TO-THE-NEW-WORLDPathological-child-voice
Speech Dataset for AI-Based Language Assessment in Children
The "Speech Database of Typically Developing and Speech-Impaired Children" is an open speech dataset designed to support the development of AI-based language assessment systems. It contains speech samples from children aged 2 to 9 who are either typically developing or have reduced consonant articulation accuracy.
This dataset is based on standardized Korean articulation tools:
APAC (Articulation and Phonology Assessment… See the full description on the dataset page: https://huggingface.co/datasets/K-Univ/Pathological-child-voice.original_pathological_speechorpheus_tts_knn_vc_pathological
Combined Orpheus TTS 3B + Same-Speaker KNN Voice Conversion Dataset
Dataset Overview
This dataset contains fine-tuned Orpheus TTS 3B synthetic speech enhanced with same-speaker KNN voice conversion a
Enhancement Method: Orpheus TTS 3B LoRA fine-tuned synthetic speech → Same-speaker KNN voice conversion using real audio references from the same speaker
Dataset Statistics
Total Samples: 759
Total Duration: 2583.38 seconds (43.06 minutes)
Speakers: 8 speakers… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/orpheus_tts_knn_vc_pathological.spark_pathological_synthetic_data
Combined Fine-tuned Spark TTS Synthetic Speech Dataset
Dataset Statistics
Total Samples: 785
Total Duration: 0.8 hours
Speakers: 8
Corpora: TORGO, UA-Speech, LibriSpeech
Audio Format: 16kHz WAV
Data Format
Each sample contains:
audio: Audio array with sampling_rate (16kHz)
text: Original transcript text
speaker_id: Speaker identifier (F04, M02, FC02, MC01, F02, M04, 211, 4014)
corpus: Source corpus (TORGO, UA-Speech, LibriSpeech)
condition: Speaker condition… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/spark_pathological_synthetic_data.femh-pathological-voicespark_tts_knn_vc_pathological
Dataset Overview
Total Samples: 785
Total Duration: 3006.46 seconds (50.11 minutes)
Speakers: 8 speakers
Corpora: TORGO, UA-Speech, LibriSpeech
Sample Rate: 16kHz (KNN-VC output rate, native Spark compatibility)
Audio Format: WAV
MELD-videos-absolute-pathspathology_trainpathology_testsesame_pathological_synthetic_data
Combined Fine-tuned Sesame CSM 1B Synthetic Speech Dataset
Data Format
Each sample contains:
audio: Audio array with sampling_rate (24kHz)
text: Original transcript text
speaker_id: Speaker identifier (F04, M02, FC02, MC01, F02, M04, 211, 4014)
corpus: Source corpus (TORGO, UA-Speech, LibriSpeech)
condition: Speaker condition (Dysarthric, Healthy)
model_name: Fine-tuned model name
model_type: "sesame_csm_1b_adapter"
base_model: Base model used (unsloth/csm-1b)… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/sesame_pathological_synthetic_data.sesame_tts_knn_vc_pathological
Dataset Overview
Total Samples: 759
Total Duration: 2892.82 seconds (48.21 minutes)
Speakers: 8 speakers
Corpora: TORGO, UA-Speech, LibriSpeech
Sample Rate: 16kHz (KNN-VC output rate)
Audio Format: WAV
pathological_knn_vc
Combined Same-Speaker KNN Voice Conversion Dataset
Dataset Statistics
Total Samples: 799
Total Duration: 0.4 hours
Speakers: 8
Corpora: TORGO, UA-Speech, LibriSpeech
Audio Format: 16kHz WAV
Baseline TTS: resproj007/baseline_orpheus_3b
Speaker Breakdown
Speaker
Name
Corpus
Condition
Gender
Samples
Duration
FC02
TORGO Healthy Female
TORGO
Healthy
Female
100
146.0s
M04
UA-Speech Male
UA-Speech
Dysarthric
Male
200
245.7s
M02
TORGO Dysarthric Male… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/pathological_knn_vc.orpheus_pathological_synthetic_data
Combined Fine-tuned Orpheus TTS 3B Synthetic Speech Dataset
Data Format
Each sample contains:
audio: Audio array with sampling_rate (24kHz)
text: Original transcript text
speaker_id: Speaker identifier (F04, M02, FC02, MC01, F02, M04, 211, 4014)
corpus: Source corpus (TORGO, UA-Speech, LibriSpeech)
condition: Speaker condition (Dysarthric, Healthy)
model_name: Fine-tuned model name
model_type: "orpheus_3b_adapter"
base_model: Base model used (unsloth/orpheus-3b-0.1-ft)… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/orpheus_pathological_synthetic_data.
