kawinduwijewardhane/nllb-sinhala-english
NLLB Sinhala-English
A fine-tuned version of facebook/nllb-200-distilled-600M for bidirectional Sinhala ↔ English machine translation, optimized for short-form UI, application, and conversational text.
Model description
This model is based on Meta's NLLB-200 Distilled 600M checkpoint and has been fine-tuned on a custom parallel corpus of Sinhala-English sentence pairs. It supports translation in both directions:
- Sinhala (
sin_Sinh) → English (eng_Latn) - English (
eng_Latn) → Sinhala (sin_Sinh)
The model is designed specifically for short text commonly found in user interfaces, applications, games, and everyday conversations.
Intended uses
This model is suitable for:
- User interface localization
- Mobile and web application text
- Chat and conversational translation
- Short messages and notifications
- General Sinhala ↔ English translation of short-form content
Limitations
This model is not intended for:
- Long-form document translation
- Literary or creative writing
- Legal, medical, or other safety-critical content
- Highly technical or scientific documents
Additional limitations:
- Training data primarily consists of short UI and application strings, so performance may decrease on formal or domain-specific text.
- The model has not been specifically evaluated for bias, fairness, or toxicity.
- Inputs longer than 64 tokens may be truncated during inference.
Training data
The model was trained on a custom Sinhala-English parallel dataset collected from UI, application, and game-related text.
Dataset preparation included:
- Removal of duplicate sentence pairs
- Removal of empty and invalid entries
- Whitespace normalization
- Bidirectional expansion (Sinhala → English and English → Sinhala)
Dataset statistics
- Parallel sentence pairs: 34,414
- Bidirectional training examples: 61,944
- Train / Validation / Test split: 90% / 5% / 5%
- Random seed: 42
Training
Base model
facebook/nllb-200-distilled-600M
Hardware
- NVIDIA Tesla T4 GPU
Training configuration
- Optimizer: Adafactor
- Learning rate: 1e-4
- Batch size: 8
- Gradient accumulation: 4
- Effective batch size: 32
- Mixed precision: FP16
- Epochs: 3
- Maximum sequence length: 64 tokens
- Gradient checkpointing: Enabled
Framework
- Hugging Face Transformers
- Seq2SeqTrainer
Evaluation
The model was evaluated on a held-out test set consisting of 1,721 sentence pairs covering both translation directions.
Evaluation metrics:
- BLEU
- chrF
Usage
Transformers
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
tokenizer = AutoTokenizer.from_pretrained(
"kawinduwijewardhane/nllb-sinhala-english"
)
model = AutoModelForSeq2SeqLM.from_pretrained(
"kawinduwijewardhane/nllb-sinhala-english"
)
tokenizer.src_lang = "sin_Sinh"
inputs = tokenizer(
"ඔබට කොහොමද",
return_tensors="pt"
)
generated = model.generate(
**inputs,
forced_bos_token_id=tokenizer.convert_tokens_to_ids("eng_Latn"),
max_length=64,
)
print(tokenizer.batch_decode(generated, skip_special_tokens=True))CPU Inference
An INT8-quantized CTranslate2 version is also available for efficient CPU deployment:
https://huggingface.co/kawinduwijewardhane/nllb-sinhala-english-ct2
License
This model inherits the license of its base model (facebook/nllb-200-distilled-600M) and is released under CC BY-NC 4.0.
Commercial use is not permitted under this license. Please review the original NLLB license before using this model in commercial applications.
Citation
If you use or reference the underlying NLLB model, please cite:
NLLB Team et al. (2022).
No Language Left Behind: Scaling Human-Centered Machine Translation.