CoolFace
Modelpublic

KasuleTrevor/qwen3-asr-lg-atypical-promptless-specaug-encoder

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes14downloads
Model Card

Qwen3-ASR Atypical Luganda — Encoder Adaptation + SpecAugment

This repository contains the validation-selected encoder checkpoint from the promptless Qwen3-ASR atypical-Luganda architectural ablation.

Experiment

  • —Base model: KasuleTrevor/cdli-qwen3-asr-lg-typical-1p7b-base-finetune
  • —Dataset: cdli/ugandan_luganda_nonstandard_speech_v1.0
  • —Condition: encoder
  • —Prompt: none
  • —SpecAugment: enabled, train-only
  • —LR: 0.0001
  • —Scheduler: cosine
  • —Seed: 42
  • —Batch / grad accumulation: 4 / 4
  • —Effective batch: 16

SpecAugment

SettingValue
Probability0.5
Time masks2
Time width40
Frequency masks2
Frequency width12

Trainable parameters

MetricValue
Total2,038,052,480
Trainable314,328,704
Trainable %15.4230%

Validation checkpoint selection

Selection was performed on the validation split only using normalized corpus WER/CER.

MetricValue
Selected checkpointcheckpoint-1500
Mean capped WER72.50%
Mean capped CER39.35%
Corpus WER100.80%
Corpus CER60.06%

Final test results

Primary reporting uses normalized per-utterance WER/CER, capped at 1.0 per utterance and then averaged. Normalized corpus metrics are retained as secondary diagnostics.

MetricResult
Mean capped utterance WER67.71%
Mean capped utterance CER31.38%
Normalized corpus WER83.54%
Normalized corpus CER44.98%
Relative primary WER reduction vs base11.69%

Results by severity

GroupnSpeakersMean capped WERMean capped CERCorpus WERCorpus CER
Mild (easily understood with minimal effort)366360.66%24.21%76.46%39.25%
Moderate (requires effort to understand)347365.30%28.10%76.28%36.91%
Severe (frequent breakdowns)315378.57%43.31%100.24%62.60%

Results by disorder

GroupnSpeakersMean capped WERMean capped CERCorpus WERCorpus CER
Articulation Disorders209268.41%30.83%91.40%51.32%
Dysarthria276260.17%25.05%65.72%29.26%
Stuttering (Disfluency Disorders)438468.45%31.18%84.72%42.44%
Voice disorder105183.07%49.92%98.77%64.39%

Reproducibility

The exact corrected trainer is archived in this repository under:

code/qwen3_asr_sft_fixed_ablation.py

The companion private results repository is:

KasuleTrevor/qwen3-asr-lg-atypical-promptless-specaug-results

No raw participant audio is redistributed.