CoolFace
Modelpublic

savinugunarathna/Small100-Singlish-Sinhala-Merged

sourceHugging Facemitupdated 7mo agoView on Hugging Face
1likes8downloads
Model Card

Small100 — Singlish → Sinhala Transliteration

Fine-tuned version of alirezamsh/small100 for the task of Singlish-to-Sinhala transliteration, developed as part of the IndoNLP 2025 Shared Task on Singlish–Sinhala Transliteration.

This is the merged (LoRA weights absorbed) final model.


Task

Singlish (romanised colloquial Sinhala) → Sinhala script transliteration.

Input (Singlish)Output (Sinhala)
mama giyaමම ගිය
kohomadaකොහොමද

Training Pipeline

Trained using a three-phase curriculum strategy with LoRA applied to all attention projection and feed-forward layers.

Data

SplitSourceSize
Phase 1 & 2 trainingphonetic_train_1M.csv1,000,000 samples
Adhoc fine-tuningadhoc.csv11,937 samples
Phonetic validationphonetic_test.csv10,003 samples
Adhoc validationadhoc_test.csv5,003 samples

Synthetic Augmentation

Adhoc data was expanded with a rule-based Singlish augmenter simulating natural romanisation variation:

  • —Vowel dropping — randomly drops non-boundary vowels (e.g. kohomada → khmada)
  • —Cluster simplification — collapses common digraphs (th→t, sh→s, nd→n, etc.)
  • —Vowel swapping — substitutes phonetically similar vowels (a↔e, i↔e, o↔u)

Aggression factor: 0.5. Applied at 15% / 20% / 15% across the three phases.

Three-Phase Curriculum

PhaseDataEpochsLRValidationAug
1 — Foundation65% of phonetic train (~650K)21e-4Phonetic15%
2 — ExpansionRemaining phonetic + 5× adhoc + 80K replay25e-5Adhoc20%
3 — Mastery10× adhoc + 200K phonetic mix22e-5Adhoc15%

Each phase resumes from the previous phase's LoRA adapter. Early stopping: patience=5, metric=CER.

LoRA Configuration

ParameterValue
Rank (r)64
Alpha128
Dropout0.05
Target modulesq_proj, k_proj, v_proj, out_proj, fc1, fc2
Trainable params19,267,584 / 352,003,072 (5.47%)

Training Arguments

ParameterValue
Batch size8
Gradient accumulation4 (effective batch: 32)
Weight decay0.01
Max grad norm1.0
Warmup ratio0.03
OptimizerAdamW fused
Precisionbfloat16 / fp16

Evaluation Results

Test SetCER ↓WER ↓BLEU ↑BERTScore ↑
Phonetic0.02110.09700.77230.9906
Adhoc0.04610.16530.64390.9899
BERTScore computed using Ransaka/sinhala-bert-medium-v2.

Small100 achieved the best average CER (0.0336) among all models evaluated on this shared task.


Usage

python
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

model_id = "savinugunarathna/Small100-Singlish-Sinhala-Merged"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)

tokenizer.src_lang = "en"
tokenizer.tgt_lang = "si"

inputs = tokenizer("mama giya", return_tensors="pt")
outputs = model.generate(**inputs, num_beams=4, max_length=128)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
# → මම ගිය