CoolFace
Datasetpublic

resproj007/pathological_speech

Pathological Speech (TORGO + UA-Speech + LibriSpeech Normal) Mixed-corpus speech dataset for training and evaluating controllable speech-synthesis and severity-classification models. Three corpora are merged with unified metadata so a single model can learn severity- and gender-conditioned generation without confounds. Splits (speaker-disjoint since 2026-09-14) Split Rows Bytes (parquet) What it is train 37704 5,500,328,387 every clip of every speaker… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/pathological_speech.

sourceHugging Faceupdated 8d agoView on Hugging Face
0likes1.2kdownloads
Dataset Card

Pathological Speech (TORGO + UA-Speech + LibriSpeech Normal)

Mixed-corpus speech dataset for training and evaluating controllable speech-synthesis and severity-classification models. Three corpora are merged with unified metadata so a single model can learn severity- and gender-conditioned generation without confounds.

Splits (speaker-disjoint since 2026-09-14)

SplitRowsBytes (parquet)What it is
train377045,500,328,387every clip of every speaker not held out
test8812930,824,798every clip of the nine held-out speakers, rows shuffled (seed 0)

No speaker appears in both splits. The held-out speakers are TORGO_F03, TORGO_FC01, TORGO_M04, TORGO_MC02, UA-Speech_M09, LibriSpeech_1263, LibriSpeech_3242, LibriSpeech_6181, LibriSpeech_6818. test is the complete held-out set; evaluations that need fewer clips subsample it with a recorded seed rather than the dataset carrying a second split.

What was wrong before

Revision 71f6861aa9ea5bcc89d35fd62d8581853c54a43c claimed a held-out test set, but its test split was a per-clip sample of the SAME speakers as train: 17 of its 21 test speakers had training clips, and all 1500 TORGO and UA-Speech test rows came from speakers with 38,741 training rows. Only the 280 LibriSpeech test rows were speaker-disjoint. Every held-out number measured on that split was measured on voices the model had trained on. Load revision="71f6861aa9ea5bcc89d35fd62d8581853c54a43c" to reproduce that split. (Revision c76c0c09f329724734e7e1b93dd13dae4db69aa8 was the first speaker-disjoint rebuild; it split the held-out clips into test and heldout, now folded into one test.)

How the held-out speakers were chosen

Per gender x pathology cell over the two dysarthric corpora, one whole speaker per corpus is held out (a corpus with a single speaker in a cell cannot be held out: UA-Speech has one female speaker, UA-Speech_F04, so that corpus x cell is untestable). The joint assignment is a deterministic search that (1) keeps every severity band of every cell in train with at least one speaker and keeps every corpus x cell in train, (2) maximises the number of severity bands present in test, (3) minimises the largest share of any band's rows that leaves training, (4) minimises the total rows held out, ties broken by speaker id. LibriSpeech's four test speakers were already disjoint and are kept. 1 exact duplicate row(s) (same speaker, same audio bytes) were dropped.

Text is NOT disjoint: TORGO and UA-Speech use fixed prompt lists, so a test speaker's words and sentences are also spoken by training speakers.

Composition

By corpus:

corpustraintest
LibriSpeech5996280
TORGO75873177
UA-Speech241215355

By corpus, gender and severity:

corpusgenderseveritytraintest
LibriSpeechfemaleNormal2224140
LibriSpeechmaleNormal3772140
TORGOfemaleModerate01086
TORGOfemaleNormal1922302
TORGOfemaleSevere2360
TORGOmaleMild8100
TORGOmaleModerate5870
TORGOmaleNormal32891123
TORGOmaleSevere743666
UA-SpeechfemaleModerate52510
UA-SpeechmaleMild53555355
UA-SpeechmaleModerate107100
UA-SpeechmaleSevere28050

By speaker:

corpusspeaker_idtraintest
LibriSpeechLibriSpeech_103660
LibriSpeechLibriSpeech_1034960
LibriSpeechLibriSpeech_1069660
LibriSpeechLibriSpeech_1088630
LibriSpeechLibriSpeech_1183370
LibriSpeechLibriSpeech_1246700
LibriSpeechLibriSpeech_1263070
LibriSpeechLibriSpeech_1363520
LibriSpeechLibriSpeech_150660
LibriSpeechLibriSpeech_1502790
LibriSpeechLibriSpeech_16241010
LibriSpeechLibriSpeech_1631120
LibriSpeechLibriSpeech_17431030
LibriSpeechLibriSpeech_1992370
LibriSpeechLibriSpeech_2011270
LibriSpeechLibriSpeech_2182640
LibriSpeechLibriSpeech_2196530
LibriSpeechLibriSpeech_226650
LibriSpeechLibriSpeech_2331100
LibriSpeechLibriSpeech_250490
LibriSpeechLibriSpeech_25141080
LibriSpeechLibriSpeech_29891000
LibriSpeechLibriSpeech_3071240
LibriSpeechLibriSpeech_3111220
LibriSpeechLibriSpeech_3112700
LibriSpeechLibriSpeech_32690
LibriSpeechLibriSpeech_32141160
LibriSpeechLibriSpeech_32401270
LibriSpeechLibriSpeech_3242071
LibriSpeechLibriSpeech_37231190
LibriSpeechLibriSpeech_3741130
LibriSpeechLibriSpeech_38301180
LibriSpeechLibriSpeech_3983680
LibriSpeechLibriSpeech_403710
LibriSpeechLibriSpeech_4214480
LibriSpeechLibriSpeech_426600
LibriSpeechLibriSpeech_4297510
LibriSpeechLibriSpeech_4362540
LibriSpeechLibriSpeech_4461110
LibriSpeechLibriSpeech_47881070
LibriSpeechLibriSpeech_4811280
LibriSpeechLibriSpeech_51041120
LibriSpeechLibriSpeech_51921190
LibriSpeechLibriSpeech_53221130
LibriSpeechLibriSpeech_53901160
LibriSpeechLibriSpeech_54561120
LibriSpeechLibriSpeech_56781080
LibriSpeechLibriSpeech_57031130
LibriSpeechLibriSpeech_57501220
LibriSpeechLibriSpeech_5778600
LibriSpeechLibriSpeech_5789570
LibriSpeechLibriSpeech_587630
LibriSpeechLibriSpeech_6181069
LibriSpeechLibriSpeech_63671130
LibriSpeechLibriSpeech_6385680
LibriSpeechLibriSpeech_6476620
LibriSpeechLibriSpeech_65291070
LibriSpeechLibriSpeech_6563920
LibriSpeechLibriSpeech_6818070
LibriSpeechLibriSpeech_7078630
LibriSpeechLibriSpeech_7113510
LibriSpeechLibriSpeech_7178650
LibriSpeechLibriSpeech_7312260
LibriSpeechLibriSpeech_73671190
LibriSpeechLibriSpeech_74471140
LibriSpeechLibriSpeech_75051150
LibriSpeechLibriSpeech_7635770
LibriSpeechLibriSpeech_7800730
LibriSpeechLibriSpeech_8014450
LibriSpeechLibriSpeech_82261180
LibriSpeechLibriSpeech_8238650
LibriSpeechLibriSpeech_8324640
LibriSpeechLibriSpeech_87701110
LibriSpeechLibriSpeech_8975530
TORGOTORGO_F012360
TORGOTORGO_F0301086
TORGOTORGO_FC010302
TORGOTORGO_FC0319220
TORGOTORGO_M017430
TORGOTORGO_M038100
TORGOTORGO_M040666
TORGOTORGO_M055870
TORGOTORGO_MC0201123
TORGOTORGO_MC0316690
TORGOTORGO_MC0416200
UA-SpeechUA-Speech_F0452510
UA-SpeechUA-Speech_M0128050
UA-SpeechUA-Speech_M0553550
UA-SpeechUA-Speech_M0753550
UA-SpeechUA-Speech_M0905355
UA-SpeechUA-Speech_M1053550

Clips with four or more words: train 7888, test 1031.

Corpora

CorpusSpeaker role
TORGOdysarthric + control
UA-Speechdysarthric (cerebral palsy)
LibriSpeechNormal supplement (train-clean-100)

Why LibriSpeech?

Without it, every severity == Normal row came from TORGO controls, meaning a model could trivially learn Normal <-> TORGO acoustics instead of true prosodic normality. The LibriSpeech train-clean-100 supplement breaks that corpus x severity confound. All LibriSpeech rows are labelled severity="Normal", condition="Control", diagnosis="None".

Metadata canonicalization (applied to all shards)

  1. 1.`gender` is always lowercase ("male" / "female").
  2. 2.`speaker_id` is prefixed with `{corpus}_` (e.g. TORGO_M01, UA-Speech_M01, LibriSpeech_3242). TORGO and UA-Speech both use M01 / M05 for different humans.
  3. 3.`severity` is recomputed per speaker from the canonical mapping (TORGO: clinical severity from the TORGO speaker sheet; UA-Speech: intelligibility banding -> Very-low (<=19%) -> Severe, Low (28-39%) / Mid (58-62%) -> Moderate, High (>=86%) -> Mild).

Other columns (diagnosis, condition, intelligibility) are preserved verbatim from their source corpora. Audio bytes are unchanged from the original revision, so content-addressed keys (sha1 of the encoded bytes) still match.

Schema

ColumnTypeNotes
audioAudio(16 kHz)WAV (PCM16, 16 kHz mono) across all corpora
textstringtranscript
speaker_idstring{corpus}_{original_id}
corpusstringTORGO, UA-Speech, or LibriSpeech
genderstringmale \female
conditionstringDysarthric, Control, ...
diagnosisstringfree text ("Cerebral palsy", "None", ...)
severitystringNormal, Mild, Moderate, Severe
intelligibilitystringcorpus-specific label (may be empty)
durationfloat64 (seconds)

Load

python
from datasets import load_dataset
ds = load_dataset("resproj007/pathological_speech")
print(ds)
# DatasetDict({
#   train: Dataset({ features: [...], num_rows: 37704 })
#   test:  Dataset({ features: [...], num_rows: 8812 })
# })

Intended use

  • Controllable pathological-speech synthesis (severity + gender conditioning).
  • Severity classification with leave-speaker-out evaluation.
  • Cross-corpus robustness analysis (TORGO <-> UA-Speech).

Licensing

Derived from corpora released under research-use terms. Redistribution is intended for academic research; verify the TORGO, UA-Speech, and LibriSpeech licenses for your use case.