datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sawit-ttudataseten-si-translation-weblate-technical-1k
En Si Translation Weblate Technical 1K
Dataset Summary
English-Sinhala Technical and UI Localization dataset with ~1,000 rows targeting software interfaces, technical terminology, and static variables.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 1000
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and extracted from the… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-weblate-technical-1k.en-si-translation-opus-conversational-4k
En Si Translation Opus Conversational 4K
Dataset Summary
English-Sinhala Conversational Translation dataset containing ~4,000 sentences capturing natural dialogue, spoken pacing, and everyday expressions, derived from the OPUS-100 corpus.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 4000
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-opus-conversational-4k.en-si-translation-cultural-idioms-500
En Si Translation Cultural Idioms 500
Dataset Summary
English-Sinhala Idioms and Cultural Expressions dataset featuring ~500 items designed to teach semantic context mapping over literal word-for-word translation transitions.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 500
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-cultural-idioms-500.en-si-translation-wmt-internet-1300
En Si Translation Wmt Internet 1300
Dataset Summary
English-Sinhala Web Forum and Internet Translation dataset containing ~1,300 high-quality sentences capturing internet slang and long-form narrative text from WMT20.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 1300
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and extracted… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-wmt-internet-1300.en-si-translation-nlpc-news-1200
En Si Translation Nlpc News 1200
Dataset Summary
English-Sinhala Formal News Translation dataset featuring ~1,200 sentences highlighting passive voice, bureaucratic syntax, and journalism-grade vocabulary.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 1200
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and extracted from the… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-nlpc-news-1200.en-si-llama3-translation-master-10k
En Si Llama3 Translation Master 10K
Dataset Summary
The complete, production-ready master alignment instruction-tuning dataset containing 10,000 highly orthogonal examples, fully mixed and wrapped directly in the official Meta Llama 3 Chat Template structural formatting.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 10000
Internal Storage Structure: Single-File data.json
Curation & Data Lineage… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-llama3-translation-master-10k.en-si-translation-flores-factual-2k
En Si Translation Flores Factual 2K
Dataset Summary
English-Sinhala Factual Translation dataset containing ~2,000 highly accurate sentences covering diverse factual domains, curated from the FLORES+ benchmark.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 2000
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and extracted from… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-flores-factual-2k.SAWiT_Hackathon_DatasetSAWIT
