CoolFace
Datasetpublic

SAWithanage/en-si-translation-opus-conversational-4k

En Si Translation Opus Conversational 4K Dataset Summary English-Sinhala Conversational Translation dataset containing ~4,000 sentences capturing natural dialogue, spoken pacing, and everyday expressions, derived from the OPUS-100 corpus. Engineering Pipeline Parameters Language Pair: English (en) to Sinhala (si) Total Valid Token Rows: 4000 Internal Storage Structure: Single-File data.json Upstream Source Attribution This specific… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-opus-conversational-4k.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes21downloads
Dataset Card

En Si Translation Opus Conversational 4K

Dataset Summary

English-Sinhala Conversational Translation dataset containing ~4,000 sentences capturing natural dialogue, spoken pacing, and everyday expressions, derived from the OPUS-100 corpus.

Engineering Pipeline Parameters

  • —Language Pair: English (en) to Sinhala (si)
  • —Total Valid Token Rows: 4000
  • —Internal Storage Structure: Single-File data.json

Upstream Source Attribution

This specific sub-split was compiled and extracted from the official open-source repository:


This repository represents a curated unit of a larger 10,000-row master training grid designed to alignment-tune 8B architectural base configurations for context-aware language tasks.