CoolFace
Modelpublic

Rekin226/nllb-3.3B-moore-lora-v0

sourceHugging Facecc-by-nc-4.0updated 2mo agoView on Hugging Face
0likes5downloads
Model Card

NLLB-200 Mooré LoRA (3.3B)

LoRA adapter for facebook/nllb-200-3.3B fine-tuned on ~205k cleaned Mooré↔French/English sentence pairs, all four directions (eng_Latn↔mos_Latn, fra_Latn↔mos_Latn).

Part of Mooré-Voice — open translation and speech recognition for Mooré (Mòoré / Mossi, ISO 639-3 mos), spoken by ~8 million people in and around Burkina Faso.

Evaluation

ModelDirectionBLEUchrF++
zero-shot baseeng_Latn→mos_Latn3.7223.77
zero-shot basefra_Latn→mos_Latn3.1822.82
zero-shot basemos_Latn→eng_Latn10.7732.01
zero-shot basemos_Latn→fra_Latn9.1429.86
fine-tunedeng_Latn→mos_Latn3.8823.41
fine-tunedfra_Latn→mos_Latn3.523.32
fine-tunedmos_Latn→eng_Latn11.8633.06
fine-tunedmos_Latn→fra_Latn11.2432.66

Training data

Curated corpus v0.1 (see repo data/CORPORA.md): MT560 (Bible-register, ~89%), community instruction pairs, NLLB-mined bitext (LASER ≥ 1.15), translatewiki. Detokenised, LID-gated, FLORES-decontaminated. FLORES-200 devtest held out for eval.

Limitations

  • —Register skew: mostly religious text → weaker on administrative/technical register.
  • —Mooré orthography follows the 1976/2003 standard as used by the source corpora; diacritic usage varies upstream.
  • —Not human-evaluated yet; BLEU/chrF++ on FLORES only.

License note

Released CC-BY-NC-4.0 because a large share of the training text derives from sources whose redistribution terms are research-use-only or undeclared (see the repo's data/CORPORA.md / data/AUDIO_CORPORA.md). A fully permissive release is planned once the corpus is rebuilt on cleared sources (Common Voice mos + translatewiki + NLLB-mined).