CoolFace
Datasetpublic

SAWithanage/en-si-parallel-3k

Dataset Card for en-si-parallel-3k Dataset Summary The en-si-parallel-3k dataset is a high-quality, synthetically generated parallel corpus containing 3,000 English-Sinhala translation pairs. It is specifically designed for fine-tuning Large Language Models (LLMs) to enhance English-to-Sinhala translation capabilities and cross-lingual understanding. Dataset Composition The dataset is structured into 60 distinct batches of 50 examples each, covering… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-parallel-3k.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
1likes21downloads

SAWithanage/en-si-parallel-3k · main · files are served by the source, never re-hosted here