CoolFace
Modelpublic

cdli/whisper-small_finetuned_ghanian_ga_nonstandard_speech_v1.0

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes7downloads
Model Card

This is a fine-tuned version of **`cdli/whisper-small_finetuned_ghanian_ga_standard_speech_v1.0`** for Ga non-standard speech. It is part of CDLI's effort to make speech technology work for people whose speech is underserved by mainstream ASR systems.

All CDLI models and datasets can be found on **CDLI's HuggingFace page**.

Dataset

The model has been fine-tuned using `cdli/ghanian_ga_nonstandard_speech_v1.0`, a dataset of speech samples of people living with impaired speech across a range of impairment severity levels and etiologies.

Note: Given the small number of speakers (especially in the test split), expect limited generalization to unseen speakers; this dataset as it is right now likely lends itself better to speaker personalization than to a general fine-tuned model for non-standard speech.

Training

The train split was used for training, and the dev split for selecting the best checkpoint.

This Whisper model was fine-tuned and is decoded using the Yoruba (yo) language setting — out of all languages Whisper supports, the one most similar to Ga.

All model parameters (encoder, decoder, and output projection) were fine-tuned.

Evaluation

This model was evaluated on the `test` split of the dataset. Utterances longer than 30 seconds were excluded:

  • —Examples evaluated: 859
  • —Speakers: 4

For decoding we ran Whisper with language=yo, task=transcribe, greedy search (num_beams=1, do_sample=False).

Results are compared against the unadapted base model `cdli/whisper-small_finetuned_ghanian_ga_standard_speech_v1.0`, evaluated identically, to show the effect of fine-tuning on non-standard speech.

We report two complementary word error rate (WER) metrics, both computed on text normalized with Whisper's BasicTextNormalizer:

  • —Standard (corpus-level) WER — the usual error rate, pooling all reference words and edit errors across the entire test set.
  • —Per-utterance averaged WER — WER computed separately for each utterance, each capped at 1.0, then averaged across utterances.

The per-utterance averaged WER bounds each utterance to [0, 1] and weights all utterances equally, so it reflects typical performance without a few catastrophic utterances dominating — but it is not a true error rate and isn't directly comparable to other published WER, hence we report the standard, corpus-level WER as well.

Results

Overall Results

ModelStandard WERPer-utterance averaged WER
Adapted0.540.51
Unadapted0.660.59
Relative improvement18%13%

Detailed Analysis

Aggregated results can hide important underlying patterns, so we also break the WER down by subset: per speaker, and — where speaker severity is available — per impairment severity group.

Results by impairment severity

All WER values below are the per-utterance averaged WER, first averaged per speaker and then averaged within each severity group. n_speakers and n_utterances are the number of speakers and test utterances in each group.

severityn_speakersn_utterancesAvg WER (unadapted model)Avg WER (adapted model)Rel. improvement
mild24430.450.3815%
moderate24160.750.6711%
Results by speaker

Per-utterance averaged WER per speaker. n_utterances is the number of test utterances for that speaker.

speaker_idseverityetiologyn_utterancesAvg WER (unadapted model)Avg WER (adapted model)Rel. improvement
123mildDevelopmental1910.410.410%
800mildCleft Palate2520.480.3527%
124moderateCerebral Palsy1840.820.759%
786moderateCleft Palate2320.680.5914%