Rekin226/whisper-small-moore-v0
Whisper-small Mooré ASR
openai/whisper-small fine-tuned for Mooré speech→text on ~38k transcribed utterances (anchor language token: yo; always pass language='yo', task='transcribe').
Part of Mooré-Voice — open translation and speech recognition for Mooré (Mòoré / Mossi, ISO 639-3 mos), spoken by ~8 million people in and around Burkina Faso.
Evaluation
Pending — run `scripts/evaluate.py`.
Training data
Community Mooré audio from Hugging Face (see repo data/AUDIO_CORPORA.md): hfdjobii TTS sets + Minervus00 collection. Audio is NOT redistributed — weights only.
Limitations
- Read speech dominates → spontaneous/telephone speech will degrade.
- Speaker diversity is limited.
- WER normalisation strips punctuation/case.
License note
Released CC-BY-NC-4.0 because a large share of the training text derives from sources whose redistribution terms are research-use-only or undeclared (see the repo's data/CORPORA.md / data/AUDIO_CORPORA.md). A fully permissive release is planned once the corpus is rebuilt on cleared sources (Common Voice mos + translatewiki + NLLB-mined).
