datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SawitMVC
SawitMVC
SawitMVC is a multi-view oil palm fruit bunch detection and counting dataset. It contains expert-reviewed YOLO annotations and per-tree JSON ground truth for counting unique fruit bunches across 4-8 camera views.
Dataset Summary
Property
Value
Trees
953 (DAMIMAS: 854, LONSUM: 99)
Images
3,992 (960 x 1280 px, JPEG)
Views per tree
4 sides (45 trees have 8 sides)
Annotation format
YOLO v8 labels + JSON ground truth
Classes
4 maturity… See the full description on the dataset page: https://huggingface.co/datasets/ULM-DS-Lab/SawitMVC.Sawit-Weight
Sawit-Weight
Two-view field photographs of oil palm fresh fruit bunches (FFB, tandan buah segar), each
paired with a bounding box and the ground-truth weight measured on a scale at the collection
point. The dataset targets vision-based weight estimation and bunch detection for smallholder
and estate harvest logistics.
Ringkasan: 31 tandan buah segar kelapa sawit varietas TANERA dari blok 303, difoto dari dua
sisi dan ditimbang langsung di lapangan. Setiap gambar disertai kotak… See the full description on the dataset page: https://huggingface.co/datasets/ULM-DS-Lab/Sawit-Weight.en-si-gemma3-translation-master-10ksawit-ttudataseten-si-translation-weblate-technical-1k
En Si Translation Weblate Technical 1K
Dataset Summary
English-Sinhala Technical and UI Localization dataset with ~1,000 rows targeting software interfaces, technical terminology, and static variables.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 1000
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and extracted from the… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-weblate-technical-1k.en-si-parallel-3k
Dataset Card for en-si-parallel-3k
Dataset Summary
The en-si-parallel-3k dataset is a high-quality, synthetically generated parallel corpus containing 3,000 English-Sinhala translation pairs. It is specifically designed for fine-tuning Large Language Models (LLMs) to enhance English-to-Sinhala translation capabilities and cross-lingual understanding.
Dataset Composition
The dataset is structured into 60 distinct batches of 50 examples each, covering… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-parallel-3k.en-si-translation-opus-conversational-4k
En Si Translation Opus Conversational 4K
Dataset Summary
English-Sinhala Conversational Translation dataset containing ~4,000 sentences capturing natural dialogue, spoken pacing, and everyday expressions, derived from the OPUS-100 corpus.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 4000
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-opus-conversational-4k.en-si-translation-cultural-idioms-500
En Si Translation Cultural Idioms 500
Dataset Summary
English-Sinhala Idioms and Cultural Expressions dataset featuring ~500 items designed to teach semantic context mapping over literal word-for-word translation transitions.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 500
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-cultural-idioms-500.en-si-translation-wmt-internet-1300
En Si Translation Wmt Internet 1300
Dataset Summary
English-Sinhala Web Forum and Internet Translation dataset containing ~1,300 high-quality sentences capturing internet slang and long-form narrative text from WMT20.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 1300
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and extracted… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-wmt-internet-1300.mt5-finetuned-sawit-dataset-v2en-si-parallel-3k-llama3-format
Dataset Card for en-si-parallel-3k-llama3-format
Overview
This is a specially formatted, instruction-ready version of the original en-si-parallel-3k dataset. It has been strictly engineered to fine-tune the SAWithanage/SinLlama-Llama-3-8B-Merged base model (and other Llama-3 architectures) for English-to-Sinhala translation.
This dataset utilizes the industry-standard ShareGPT Conversational Format. By structuring the data as a standardized list of roles and contents, it… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-parallel-3k-llama3-format.dataset_sawitmt5-finetuned-sawit-datasetSAWiT-Tamil-Colloquial-DatasetSawitHackathonDatasetSubmissionen-si-translation-nlpc-news-1200
En Si Translation Nlpc News 1200
Dataset Summary
English-Sinhala Formal News Translation dataset featuring ~1,200 sentences highlighting passive voice, bureaucratic syntax, and journalism-grade vocabulary.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 1200
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and extracted from the… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-nlpc-news-1200.en-si-llama3-translation-master-10k
En Si Llama3 Translation Master 10K
Dataset Summary
The complete, production-ready master alignment instruction-tuning dataset containing 10,000 highly orthogonal examples, fully mixed and wrapped directly in the official Meta Llama 3 Chat Template structural formatting.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 10000
Internal Storage Structure: Single-File data.json
Curation & Data Lineage… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-llama3-translation-master-10k.sawit-dataseten-si-parallel-3k-gemmaen-si-translation-flores-factual-2k
En Si Translation Flores Factual 2K
Dataset Summary
English-Sinhala Factual Translation dataset containing ~2,000 highly accurate sentences covering diverse factual domains, curated from the FLORES+ benchmark.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 2000
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and extracted from… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-flores-factual-2k.sawit-tamil-datasetsawit-fine-tuning-datasetsawit-swnsawit-swn1702sawit-swn1702newsawit-etsawit-etnewmt5-sawit-dataset-latestSAWiT_Hackathon_DatasetSAWIT
