savinugunarathna/Small100-Singlish-Sinhala-Merged
Small100 — Singlish → Sinhala Transliteration
Fine-tuned version of alirezamsh/small100 for the task of Singlish-to-Sinhala transliteration, developed as part of the IndoNLP 2025 Shared Task on Singlish–Sinhala Transliteration.
This is the merged (LoRA weights absorbed) final model.
Task
Singlish (romanised colloquial Sinhala) → Sinhala script transliteration.
Training Pipeline
Trained using a three-phase curriculum strategy with LoRA applied to all attention projection and feed-forward layers.
Data
Synthetic Augmentation
Adhoc data was expanded with a rule-based Singlish augmenter simulating natural romanisation variation:
- Vowel dropping — randomly drops non-boundary vowels (e.g.
kohomada→khmada) - Cluster simplification — collapses common digraphs (
th→t,sh→s,nd→n, etc.) - Vowel swapping — substitutes phonetically similar vowels (
a↔e,i↔e,o↔u)
Aggression factor: 0.5. Applied at 15% / 20% / 15% across the three phases.
Three-Phase Curriculum
Each phase resumes from the previous phase's LoRA adapter. Early stopping: patience=5, metric=CER.
LoRA Configuration
Training Arguments
Evaluation Results
BERTScore computed using Ransaka/sinhala-bert-medium-v2.
Small100 achieved the best average CER (0.0336) among all models evaluated on this shared task.
Usage
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
model_id = "savinugunarathna/Small100-Singlish-Sinhala-Merged"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)
tokenizer.src_lang = "en"
tokenizer.tgt_lang = "si"
inputs = tokenizer("mama giya", return_tensors="pt")
outputs = model.generate(**inputs, num_beams=4, max_length=128)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
# → මම ගිය