CoolFace
Modelpublic

KnoxDevelopers/English-to-Maasai-language-translation-LoRA

sourceHugging Facecc-by-nc-4.0updated 2mo agoView on Hugging Face
0likes21downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

English-to-Maasai-language-translation-LoRA

This model is a fine-tuned version of facebook/nllb-200-distilled-600M on the English-Maasai dataset. It achieves the following results on the evaluation set:

  • —Loss: 2.3882

Model description

A LoRA adapter for translating text from English to Maasai.

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 0.0002
  • —trainbatchsize: 8
  • —evalbatchsize: 8
  • —seed: 42
  • —gradientaccumulationsteps: 2
  • —totaltrainbatch_size: 16
  • —optimizer: Use OptimizerNames.ADAMWTORCHFUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • —lrschedulertype: linear
  • —num_epochs: 7
  • —mixedprecisiontraining: Native AMP

Training results

Training LossEpochStepValidation Loss
7.30631.017443.3755
6.43222.034882.9127
5.86813.052322.6750
5.48324.069762.5355
5.31535.087202.4531
5.22786.0104642.4057
5.24097.0122082.3882

Evaluation metrics

MetricResultsExplained
sacrebleu7.230Evaluation matches exact word sequence referencing our ground truth and adapter output.
0-20= translation is literal or broken, 20-50=good, >60=very good or identical training and test data.
chrF++36.480Eval matches character sequence instead of whole words.
Model adapter is getting the root words and grammar concepts right but using different word variations or spellings than our ground truth, this is concluded due to the low sacrebleu results.

Framework versions

  • —PEFT 0.19.1
  • —Transformers 5.0.0
  • —Pytorch 2.10.0+cu128
  • —Datasets 5.0.0
  • —Tokenizers 0.22.2