CoolFace
Modelpublic

KasuleTrevor/qwen3-asr-lg-atypical-promptless-specaug-decoder

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes8downloads
Model Card

Qwen3-ASR Atypical Luganda — Decoder Adaptation + SpecAugment

This repository contains the validation-selected decoder checkpoint from the promptless Qwen3-ASR atypical-Luganda architectural ablation.

Experiment

  • —Base model: KasuleTrevor/cdli-qwen3-asr-lg-typical-1p7b-base-finetune
  • —Dataset: cdli/ugandan_luganda_nonstandard_speech_v1.0
  • —Condition: decoder
  • —Prompt: none
  • —SpecAugment: enabled, train-only
  • —LR: 0.0001
  • —Scheduler: cosine
  • —Seed: 42
  • —Batch / grad accumulation: 4 / 4
  • —Effective batch: 16

SpecAugment

SettingValue
Probability0.5
Time masks2
Time width40
Frequency masks2
Frequency width12

Trainable parameters

MetricValue
Total2,038,052,480
Trainable1,720,574,976
Trainable %84.4225%

Validation checkpoint selection

Selection was performed on the validation split only using normalized corpus WER/CER.

MetricValue
Selected checkpointcheckpoint-500
Mean capped WER65.25%
Mean capped CER33.30%
Corpus WER73.59%
Corpus CER35.96%

Final test results

Primary reporting uses normalized per-utterance WER/CER, capped at 1.0 per utterance and then averaged. Normalized corpus metrics are retained as secondary diagnostics.

MetricResult
Mean capped utterance WER59.81%
Mean capped utterance CER25.88%
Normalized corpus WER76.43%
Normalized corpus CER38.76%
Relative primary WER reduction vs base21.99%

Results by severity

GroupnSpeakersMean capped WERMean capped CERCorpus WERCorpus CER
Mild (easily understood with minimal effort)366354.06%20.86%67.69%30.44%
Moderate (requires effort to understand)347358.11%23.08%72.96%33.45%
Severe (frequent breakdowns)315368.36%34.79%90.72%56.10%

Results by disorder

GroupnSpeakersMean capped WERMean capped CERCorpus WERCorpus CER
Articulation Disorders209259.95%21.87%76.37%35.15%
Dysarthria276251.19%19.24%56.13%21.92%
Stuttering (Disfluency Disorders)438461.40%27.75%77.83%40.45%
Voice disorder105175.53%43.50%110.60%72.83%

Reproducibility

The exact corrected trainer is archived in this repository under:

code/qwen3_asr_sft_fixed_ablation.py

The companion private results repository is:

KasuleTrevor/qwen3-asr-lg-atypical-promptless-specaug-results

No raw participant audio is redistributed.