SAWithanage/en-si-llama3-translation-master-10k
En Si Llama3 Translation Master 10K Dataset Summary The complete, production-ready master alignment instruction-tuning dataset containing 10,000 highly orthogonal examples, fully mixed and wrapped directly in the official Meta Llama 3 Chat Template structural formatting. Engineering Pipeline Parameters Language Pair: English (en) to Sinhala (si) Total Valid Token Rows: 10000 Internal Storage Structure: Single-File data.json Curation &… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-llama3-translation-master-10k.
En Si Llama3 Translation Master 10K
Dataset Summary
The complete, production-ready master alignment instruction-tuning dataset containing 10,000 highly orthogonal examples, fully mixed and wrapped directly in the official Meta Llama 3 Chat Template structural formatting.
Engineering Pipeline Parameters
- Language Pair: English (
en) to Sinhala (si) - Total Valid Token Rows: 10000
- Internal Storage Structure: Single-File
data.json
Curation & Data Lineage
This master dataset is a balanced compile of multiple high-quality upstream data pipelines. It contains components explicitly drawn from:
- Factual: openlanguagedata/flores_plus (2,000 examples)
- Conversational: Helsinki-NLP/opus-100 (4,000 examples)
- Formal News: NLPC-UOM/sentence_alignment_dataset-Sinhala-Tamil-English (1,200 examples)
- Technical/UI: ayymen/Weblate-Translations (1,000 examples)
- Idioms: Venuraa/English-Sinhala-Idioms-Parallel-Translations (500 examples)
- Web/Internet Forums: wmt/wmt20_mlqe_task1 (1,300 examples)
This repository represents a curated unit of a larger 10,000-row master training grid designed to alignment-tune 8B architectural base configurations for context-aware language tasks.
