CoolFace
Datasetpublic

SAWithanage/en-si-llama3-translation-master-10k

En Si Llama3 Translation Master 10K Dataset Summary The complete, production-ready master alignment instruction-tuning dataset containing 10,000 highly orthogonal examples, fully mixed and wrapped directly in the official Meta Llama 3 Chat Template structural formatting. Engineering Pipeline Parameters Language Pair: English (en) to Sinhala (si) Total Valid Token Rows: 10000 Internal Storage Structure: Single-File data.json Curation &… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-llama3-translation-master-10k.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes8downloads
Dataset Card

En Si Llama3 Translation Master 10K

Dataset Summary

The complete, production-ready master alignment instruction-tuning dataset containing 10,000 highly orthogonal examples, fully mixed and wrapped directly in the official Meta Llama 3 Chat Template structural formatting.

Engineering Pipeline Parameters

  • —Language Pair: English (en) to Sinhala (si)
  • —Total Valid Token Rows: 10000
  • —Internal Storage Structure: Single-File data.json

Curation & Data Lineage

This master dataset is a balanced compile of multiple high-quality upstream data pipelines. It contains components explicitly drawn from:

  1. 1.Factual: openlanguagedata/flores_plus (2,000 examples)
  2. 2.Conversational: Helsinki-NLP/opus-100 (4,000 examples)
  3. 3.Formal News: NLPC-UOM/sentence_alignment_dataset-Sinhala-Tamil-English (1,200 examples)
  4. 4.Technical/UI: ayymen/Weblate-Translations (1,000 examples)
  5. 5.Idioms: Venuraa/English-Sinhala-Idioms-Parallel-Translations (500 examples)
  6. 6.Web/Internet Forums: wmt/wmt20_mlqe_task1 (1,300 examples)

This repository represents a curated unit of a larger 10,000-row master training grid designed to alignment-tune 8B architectural base configurations for context-aware language tasks.