Rekin226/nllb-600M-moore-lora-v0
NLLB-200 Mooré LoRA (600M)
LoRA adapter for facebook/nllb-200-distilled-600M fine-tuned on ~205k cleaned Mooré↔French/English sentence pairs, all four directions (eng_Latn↔mos_Latn, fra_Latn↔mos_Latn).
Part of Mooré-Voice — open translation and speech recognition for Mooré (Mòoré / Mossi, ISO 639-3 mos), spoken by ~8 million people in and around Burkina Faso.
Evaluation
Training data
Curated corpus v0.1 (see repo data/CORPORA.md): MT560 (Bible-register, ~89%), community instruction pairs, NLLB-mined bitext (LASER ≥ 1.15), translatewiki. Detokenised, LID-gated, FLORES-decontaminated. FLORES-200 devtest held out for eval.
Limitations
- Register skew: mostly religious text → weaker on administrative/technical register.
- Mooré orthography follows the 1976/2003 standard as used by the source corpora; diacritic usage varies upstream.
- Not human-evaluated yet; BLEU/chrF++ on FLORES only.
License note
Released CC-BY-NC-4.0 because a large share of the training text derives from sources whose redistribution terms are research-use-only or undeclared (see the repo's data/CORPORA.md / data/AUDIO_CORPORA.md). A fully permissive release is planned once the corpus is rebuilt on cleared sources (Common Voice mos + translatewiki + NLLB-mined).
