CoolFace
Modelpublic

KasuleTrevor/cdli-qwen3-asr-lg-atypical-stage3-1p7b-base

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes12downloads
Model Card

CDLI Qwen3-ASR Luganda Atypical Speech Fine-tune

This repo contains the selected checkpoint checkpoint-500 from the LG-QWEN3-ASR-ATYPICAL-STAGE3-1P7B-BASE run.

Training Setup

  • —Base model: KasuleTrevor/cdli-qwen3-asr-lg-typical-1p7b-base-finetune
  • —Dataset: cdli/ugandanlugandanonstandardspeechv1.0
  • —Split files: train_cleaned.tsv, validation_cleaned.tsv, test_cleaned.tsv
  • —Training language tag: Luganda
  • —Forced inference language: disabled
  • —Epochs: 3
  • —Batch size: 4
  • —Gradient accumulation: 4
  • —Learning rate: 5e-05
  • —Scheduler: cosine
  • —Warmup ratio: 0.03
  • —Save steps: 500
  • —Selected checkpoint: checkpoint-500
  • —Selection reason: best validation avgwernormalized_capped (0.557665, corpus WER 0.659176)

Final Test Metrics

  • —Primary WER: avg utterance WER normalized, capped at 1.0 = 0.427355
  • —Primary CER: avg utterance CER normalized, capped at 1.0 = 0.182363
  • —Corpus WER (normalized, uncapped): 0.542166
  • —Corpus CER (normalized, uncapped): 0.247597

Checkpoint Selection Evidence

checkpointwer_normalizedcer_normalizedavg_wer_normalizedavg_cer_normalizedeval_loss
checkpoint-5000.55766491480665380.26916296176281260.55766491480665380.26916296176281260.4462866187095642
checkpoint-10000.55819210386839630.27461295582937130.55819210386839630.27461295582937130.47903984785079956
checkpoint-11070.56033341574028270.274848949931701340.56033341574028270.274848949931701340.47903984785079956

Severity Breakdown

severity_speech_impairmentn_samplesn_speakersmean_wermean_cermedian_wermedian_cer
Severe (frequent breakdowns)31530.55070.27870.56250.25
Moderate (requires effort to understand)34730.41280.16370.3750.1146
Mild (easily understood with minimal effort)36530.33480.1170.28570.0714

Artifacts

  • —Result folder: results/checkpoint-500/
  • —Includes checkpoint validation summaries, final test predictions, scored outputs, and grouped metadata analyses.