datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pathological_speech
Pathological Speech (TORGO + UA-Speech + LibriSpeech Normal)
Mixed-corpus speech dataset for training and evaluating controllable
speech-synthesis and severity-classification models. Three corpora are merged
with unified metadata so a single model can learn severity- and
gender-conditioned generation without confounds.
Splits (speaker-disjoint since 2026-09-14)
Split
Rows
Bytes (parquet)
What it is
train
37704
5,500,328,387
every clip of every speaker… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/pathological_speech.original_pathological_speech
