CoolFace
Datasetpublic

SAWithanage/en-si-parallel-3k

Dataset Card for en-si-parallel-3k Dataset Summary The en-si-parallel-3k dataset is a high-quality, synthetically generated parallel corpus containing 3,000 English-Sinhala translation pairs. It is specifically designed for fine-tuning Large Language Models (LLMs) to enhance English-to-Sinhala translation capabilities and cross-lingual understanding. Dataset Composition The dataset is structured into 60 distinct batches of 50 examples each, covering… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-parallel-3k.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
1likes21downloads
settings

This repository belongs to SAWithanage on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameen-si-parallel-3k
visibilitypublic
licencecc-by-4.0
gatedno
ownerSAWithanage
Account settings
SAWithanage/en-si-parallel-3k · CoolFace