CoolFace
Modelpublic

PrinceAlhassanNasamu/tekyerema-asr-ctc-kus-v2

sourceHugging Facecc-by-nc-4.0updated 5d agoView on Hugging Face
0likes217downloads
Model Card

Kusaal ASR (w2v-BERT CTC) — a new line, not yet promoted

Author: Prince Nasamu Alhassan

Overview

A new Kusaal line rather than a continuation, for a structural reason: the current Kusaal champion (`tekyerema-asr-mms-kus`) is an MMS adapter, and the w2v-BERT CTC trainer cannot continue one.

The base was chosen on evidence. MMS-1b-all never saw Kusaal at all, while `KhayaAI/w2v-bert-kus` is natively Kusaal-trained and scores 37.9 WER / 17.3 CER on Nsanku.

Use it

python
import torch, soundfile as sf
from transformers import AutoProcessor, AutoModelForCTC

repo  = "PrinceAlhassanNasamu/tekyerema-asr-ctc-kus-v2"
proc  = AutoProcessor.from_pretrained(repo)
model = AutoModelForCTC.from_pretrained(repo).eval()

wav, sr = sf.read("clip.wav", dtype="float32")   # 16 kHz mono
inp = proc(wav, sampling_rate=16_000, return_tensors="pt")
with torch.no_grad():
    logits = model(**inp).logits
print(proc.batch_decode(logits.argmax(-1))[0])

Status: MEASURED 2026-09-23 — and NOT promoted

This model existed for months with no benchmark at all, because the head-to-head it was built for was never run. It has now been run, on kusaal_scripture, n=100, same harness and same slice for both models:

modelWERCERs/item
this model (tekyerema-asr-ctc-kus-v2)24.619.650.365
tekyerema-asr-mms-kus (the champion)25.509.520.089
kasa42-asr (previous champion)41.0419.196.564
facebook/mms-1b-all (the base)59.3422.100.088
KhayaAI/w2v-bert-...kus... (this model's base)68.1326.14—

Read that as a tie, not a win. 0.89 WER points at n=100 is inside noise, CER moves the other way, and this model is 4.1x slower per item. On the A20e-class phone this project targets, latency is the axis that decides. A model that is statistically tied and four times slower is not an improvement, so tekyerema-asr-mms-kus remains the Kusaal recogniser to use.

What the comparison does establish. Starting from KhayaAI/w2v-bert-kus was the right call: the base scores 68.13 and this fine-tune reaches 24.61, a 43.5-point move, and it lands level with an MMS adapter that had a large head start. The w2v-BERT line is viable for Kusaal; it simply is not yet better.

Do not quote these absolutes as benchmark quality. kusaal_scripture shares narrators and recordings with training material and is held out by CLIP, not by SPEAKER. Both models are inflated by it, equally, which is why the DELTA between them is usable and the absolute values are not. The related public bible_Kusaal figure of 11.69 is invalid for the same reason and must never be cited.

Never quote a Kusaal scripture score as a benchmark

The public bible_Kusaal figure of 11.69 is INVALID. It shares narrators and recordings with the training material and was held out by CLIP rather than by SPEAKER. Roughly 8 hours of unique Kusaal audio exists in total; three of the corpora that appear distinct are the same recordings re-chunked. Any Kusaal number measured without speaker-level disjointness is measuring memory.

Training data

~6,700 rows: 3,000 kusaal_scripture + 3,000 Ghana Speech + the owner's 70 demo-exact recordings over-sampled 10×. Those 70 clips are the only spontaneous mobile-money Kusaal speech that exists anywhere, which is why they are weighted up; the over-sampling is disclosed here because the same technique, at 20×, is what contaminated the Twi whisper-tiny model.

Recipe: LR 2e-5, 3 epochs, batch 4 × grad-accum 2, 20 s clip budget (over-long rows dropped, not truncated).

epochstepvalidation loss
18160.8608
216320.8164
324480.8304

The best checkpoint is epoch 2, and what is published is epoch 3. Loss rose on the final epoch. load_best_model_at_end was not set on this run.

Licence correction

This repo declared apache-2.0, inherited from the KhayaAI base. The base is Apache; the weights are not, because they derive from non-commercial scripture audio (see `kusaal-asr-dataset`, whose rights are held by Davar, FCBH and GRN) and from Ghana Speech (CC BY-NC 4.0). Corrected to CC BY-NC 4.0.

Limitations, stated plainly

  • —A validation loss is not a benchmark. Several cards on this account used to print the training run's own eval loss under a heading that read like a result. Where this card gives a WER or CER it names the slice and the scorer; where no such number exists it says the model is unscored.
  • —Scripture-derived audio does not generalise to speech. The Kusaal and much of the Ewe material is read scripture with a small number of narrators. Held out by CLIP rather than by SPEAKER it produces scores that look excellent and mean nothing -- the public bible_Kusaal figure of 11.69 is invalid for exactly this reason and must never be quoted.
  • —Evaluation is on read and machine-translated text. No corpus of people speaking natural phone commands in these languages exists. Numbers measured this way are optimistic about phrasing and pessimistic about code-switching, and should not be read as field performance.
  • —Research work from a hackathon entry, not a supported product.

The rest of the family

Twi — `tekyerema-asr-ctc` (v1, w2v-BERT CTC) · `tekyerema-asr-ctc-twi-v2` (continuation)

Kusaal — `tekyerema-asr-mms-kus` (MMS adapter, the one to use) · `tekyerema-asr-ctc-kus-v2` (new w2v-BERT line) · `kasa42-asr` (third-party baseline)

Ewe — `tekyerema-asr-mms-ewe` (the one to use) · `tekyerema-asr-ctc-ewe` · `tekyerema-asr-ctc-ewe-v3`

Other recognisers — `tekyerema-asr-mms-dag` · `tekyerema-asr-mms-hau` · `whisper-large-v3-turbo-tekyerema-eng`

Voices — `tekyerema-tts-twi` · `tekyerema-tts-kus` · `tekyerema-tts-ewe` · `tekyerema-tts-hau` · `tekyerema-tts-eng`

Agent models — `tekyerema-1-v2-tool-pilot` · `tekyerema-1-native-v2-tool-pilot` · `tekyerema-audio-native-4k`

Translation & routing — `tekyerema-nllb600m-v1` · `kusaal-nllb-600M` · `tekyerema-intent-afroxlmr`

Acknowledgements

Compute resources provided by AI Skills and Compute Africa (AISCA). Trained on the Ghana NLP H200 GPU and on Kaggle T4s. Please keep derivatives non-commercial and share improvements back with the Ghana NLP community (ghananlpcommunity).