sinhala-nlp/Qwen3.5-27B-PaliSinhala-Pali2Si-en
Qwen3.5-27B-PaliSinhala-Pali2Si-en
A LoRA adapter for Pali to Sinhala translation, fine-tuned from Qwen/Qwen3.5-27B as part of the SinGen Sinhala text generation benchmark.
Usage
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3.5-27B")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-27B", dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, "sinhala-nlp/Qwen3.5-27B-PaliSinhala-Pali2Si-en")Prompts use the base model's chat template with thinking disabled, and the assistant response begins with the Translation: prefix.
Training
Evaluation
Last 1000 rows of sinhala-nlp/pali-sinhala, whitespace-tokenized (sacreBLEU's default 13a tokenizer splits Sinhala conjuncts and vowel signs):
Read the scores with the length ratio. The corpus is in canonical order, so the trailing 1000 rows used as the test set are much longer than the training pool (median target length differs by roughly an order of magnitude). Predictions are therefore shorter than references and BLEU is partly driven by the brevity penalty rather than translation quality. This split is kept as-is so the numbers stay comparable across the model families evaluated in SinGen.
Licence
This adapter inherits the licence of the base model; check the base model card before redistributing. Verify the terms of the sinhala-nlp/pali-sinhala dataset on its dataset card as well.
