PrinceAlhassanNasamu/whisper-large-v3-turbo-tekyerema-eng-foundation
Ghanaian English ASR — course 1 (foundation)
Author: Prince Nasamu Alhassan
Overview
The intermediate checkpoint of a two-course English run, not the finished model. Course 1 adapts whisper-large-v3-turbo on broadcast and silver-quality Ghanaian speech; course 2 then fine-tunes on the gold slice and produces whisper-large-v3-turbo-tekyerema-eng, which is the one to use.
It is published because it is expensive — roughly 12 h 48 m of streamed training — and because it carries a _course_complete.json marker that lets a rerun skip straight to course 2. Deleting it would mean paying for course 1 again.
Use it
import torch, soundfile as sf
from transformers import AutoProcessor, WhisperForConditionalGeneration
repo = "PrinceAlhassanNasamu/whisper-large-v3-turbo-tekyerema-eng-foundation"
proc = AutoProcessor.from_pretrained(repo)
# fp32 on purpose. The checkpoint is stored fp16 and audio features are
# fp32, so the first conv otherwise fails with "Input type (float) and bias
# type (c10::Half) should be the same" — and CPUs cannot do fp16 conv
# usefully anyway.
model = WhisperForConditionalGeneration.from_pretrained(
repo, dtype=torch.float32).eval()
wav, sr = sf.read("clip.wav", dtype="float32") # 16 kHz mono
feats = proc(wav, sampling_rate=16_000, return_tensors="pt").input_features
with torch.no_grad():
ids = model.generate(feats, language="en", task="transcribe")
print(proc.batch_decode(ids, skip_special_tokens=True)[0].strip())Training data
Trained on the Ghana Speech dataset and related Ghanaian corpora, licensed CC BY-NC 4.0.
Intended use & license
Non-commercial use only (CC BY-NC 4.0). This is inherited from the training data and required by the terms under which the compute was granted: models trained in that window are non-commercial by condition of access, not by inference.
Limitations, stated plainly
- Dagbani did get a recogniser, and the claim that it could not was wrong twice over. Every card on this account used to say that "one fine-tuning session on 74 validation rows would not change that". Those 74 rows are the eng-dag machine-translation validation split; the Dagbani speech data in this same account is
waxal_dag— 13,228 training rows, 1,750 validation rows, ~71 hours, 1,041 speakers with the largest at 1%. Trained on it,tekyerema-asr-mms-dagscores 36.94 / 11.71, against the 86.59 / 33.95 this project had believed was the ceiling. It still loses toFarmerlineML/w2v-bert-2.0_2026_dagbani_ASRat 29.20 / 9.27, which is what the agent actually serves. A number carried across from a translation table into a speech claim was then repeated on every card here until 2026-09-22. - Evaluation is on read and machine-translated text. No recordings of people speaking agent commands in these languages exist. Numbers measured this way are optimistic about phrasing and pessimistic about code-switching, and should not be read as field performance.
- Research work from a hackathon entry, not a supported product.
The rest of the family
Recognisers
- `whisper-large-v3-turbo-tekyerema-eng-foundation` — Ghanaian English ASR — course 1 (foundation)
- `kusaal-whisper-small-lora` — Kusaal ASR (Whisper-small LoRA, superseded)
- `kasa42-asr` — KASA-42 (Kusaal, third-party export)
- `tekyerema-asr-ctc` — Twi ASR (w2v-BERT CTC)
- `tekyerema-asr-mms-ewe` — Ewe ASR (MMS adapter)
- `tekyerema-asr-mms-dag` — Dagbani ASR (MMS adapter)
- `tekyerema-asr-mms-hau` — Hausa ASR (MMS adapter)
- `tekyerema-asr-mms-kus` — Kusaal ASR (MMS adapter)
- `whisper-large-v3-turbo-tekyerema-eng` — Ghanaian English ASR (Whisper large-v3-turbo)
Voices
- `tekyerema-tts-twi` — Twi TTS (VITS)
- `tekyerema-tts-kus` — Kusaal TTS (VITS)
- `tekyerema-tts-ewe` — Ewe TTS (VITS)
- `tekyerema-tts-hau` — Hausa TTS (VITS)
- `tekyerema-tts-eng` — Ghanaian English TTS (VITS)
Agent models
- `tekyerema-1-reply` — Tɛkyerɛma-1 reply adapter (arm ①)
- `tekyerema-1-native-reply` — Tɛkyerɛma-1 reply adapter (arm ②)
- `tekyerema-1-tool` — Tɛkyerɛma-1 tool adapter (arm 1)
- `tekyerema-audio-native` — Tɛkyerɛma-1 audio-native (arm 3)
- `tekyerema-audio-native-4k` — Tɛkyerɛma-1 audio-native, 4,000 clips (arm 3 v2)
- `tekyerema-1-native-tool` — Tɛkyerɛma-1 tool adapter (arm 2)
Translation
- `tekyerema-nllb600m-v1` — Tɛkyerɛma MT v1 (NLLB-600M)
- `kusaal-nllb-600M` — Kusaal MT specialist (NLLB-600M)
Routing
- `tekyerema-intent-afroxlmr` — Intent classifier (AfroXLMR)
Acknowledgements
Compute resources provided by AI Skills and Compute Africa (AISCA). Trained on the Ghana NLP H200 GPU. Please keep derivatives non-commercial and share improvements back with the Ghana NLP community (ghananlpcommunity).
