datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
en-si-gemma3-translation-master-10ken-si-parallel-3k
Dataset Card for en-si-parallel-3k
Dataset Summary
The en-si-parallel-3k dataset is a high-quality, synthetically generated parallel corpus containing 3,000 English-Sinhala translation pairs. It is specifically designed for fine-tuning Large Language Models (LLMs) to enhance English-to-Sinhala translation capabilities and cross-lingual understanding.
Dataset Composition
The dataset is structured into 60 distinct batches of 50 examples each, covering… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-parallel-3k.mt5-finetuned-sawit-dataset-v2en-si-parallel-3k-llama3-format
Dataset Card for en-si-parallel-3k-llama3-format
Overview
This is a specially formatted, instruction-ready version of the original en-si-parallel-3k dataset. It has been strictly engineered to fine-tune the SAWithanage/SinLlama-Llama-3-8B-Merged base model (and other Llama-3 architectures) for English-to-Sinhala translation.
This dataset utilizes the industry-standard ShareGPT Conversational Format. By structuring the data as a standardized list of roles and contents, it… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-parallel-3k-llama3-format.mt5-finetuned-sawit-datasetsawit-dataseten-si-parallel-3k-gemmasawit-tamil-datasetsawit-fine-tuning-datasetsawit-swnsawit-swn1702sawit-swn1702newsawit-etsawit-etnewmt5-sawit-dataset-latest
