CoolFace
Modelpublic

kawinduwijewardhane/nllb-sinhala-english

sourceHugging Facecc-by-nc-4.0updated 2mo agoView on Hugging Face
1likes114downloads
Model Card

NLLB Sinhala-English

A fine-tuned version of facebook/nllb-200-distilled-600M for bidirectional Sinhala ↔ English machine translation, optimized for short-form UI, application, and conversational text.

Model description

This model is based on Meta's NLLB-200 Distilled 600M checkpoint and has been fine-tuned on a custom parallel corpus of Sinhala-English sentence pairs. It supports translation in both directions:

  • —Sinhala (sin_Sinh) → English (eng_Latn)
  • —English (eng_Latn) → Sinhala (sin_Sinh)

The model is designed specifically for short text commonly found in user interfaces, applications, games, and everyday conversations.

Intended uses

This model is suitable for:

  • —User interface localization
  • —Mobile and web application text
  • —Chat and conversational translation
  • —Short messages and notifications
  • —General Sinhala ↔ English translation of short-form content

Limitations

This model is not intended for:

  • —Long-form document translation
  • —Literary or creative writing
  • —Legal, medical, or other safety-critical content
  • —Highly technical or scientific documents

Additional limitations:

  • —Training data primarily consists of short UI and application strings, so performance may decrease on formal or domain-specific text.
  • —The model has not been specifically evaluated for bias, fairness, or toxicity.
  • —Inputs longer than 64 tokens may be truncated during inference.

Training data

The model was trained on a custom Sinhala-English parallel dataset collected from UI, application, and game-related text.

Dataset preparation included:

  • —Removal of duplicate sentence pairs
  • —Removal of empty and invalid entries
  • —Whitespace normalization
  • —Bidirectional expansion (Sinhala → English and English → Sinhala)

Dataset statistics

  • —Parallel sentence pairs: 34,414
  • —Bidirectional training examples: 61,944
  • —Train / Validation / Test split: 90% / 5% / 5%
  • —Random seed: 42

Training

Base model

  • —facebook/nllb-200-distilled-600M

Hardware

  • —NVIDIA Tesla T4 GPU

Training configuration

  • —Optimizer: Adafactor
  • —Learning rate: 1e-4
  • —Batch size: 8
  • —Gradient accumulation: 4
  • —Effective batch size: 32
  • —Mixed precision: FP16
  • —Epochs: 3
  • —Maximum sequence length: 64 tokens
  • —Gradient checkpointing: Enabled

Framework

  • —Hugging Face Transformers
  • —Seq2SeqTrainer

Evaluation

The model was evaluated on a held-out test set consisting of 1,721 sentence pairs covering both translation directions.

Evaluation metrics:

  • —BLEU
  • —chrF

Usage

Transformers

python
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

tokenizer = AutoTokenizer.from_pretrained(
    "kawinduwijewardhane/nllb-sinhala-english"
)

model = AutoModelForSeq2SeqLM.from_pretrained(
    "kawinduwijewardhane/nllb-sinhala-english"
)

tokenizer.src_lang = "sin_Sinh"

inputs = tokenizer(
    "ඔබට කොහොමද",
    return_tensors="pt"
)

generated = model.generate(
    **inputs,
    forced_bos_token_id=tokenizer.convert_tokens_to_ids("eng_Latn"),
    max_length=64,
)

print(tokenizer.batch_decode(generated, skip_special_tokens=True))

CPU Inference

An INT8-quantized CTranslate2 version is also available for efficient CPU deployment:

https://huggingface.co/kawinduwijewardhane/nllb-sinhala-english-ct2

License

This model inherits the license of its base model (facebook/nllb-200-distilled-600M) and is released under CC BY-NC 4.0.

Commercial use is not permitted under this license. Please review the original NLLB license before using this model in commercial applications.

Citation

If you use or reference the underlying NLLB model, please cite:

text
NLLB Team et al. (2022).
No Language Left Behind: Scaling Human-Centered Machine Translation.