SAWit
Datasets
All datasets matching “SAWit”SawitMVC
SawitMVC
SawitMVC is a multi-view oil palm fruit bunch detection and counting dataset. It contains expert-reviewed YOLO annotations and per-tree JSON ground truth for counting unique fruit bunches across 4-8 camera views.
Dataset Summary
Property
Value
Trees
953 (DAMIMAS: 854, LONSUM: 99)
Images
3,992 (960 x 1280 px, JPEG)
Views per tree
4 sides (45 trees have 8 sides)
Annotation format
YOLO v8 labels + JSON ground truth
Classes
4 maturity… See the full description on the dataset page: https://huggingface.co/datasets/ULM-DS-Lab/SawitMVC.Sawit-Weight
Sawit-Weight
Two-view field photographs of oil palm fresh fruit bunches (FFB, tandan buah segar), each
paired with a bounding box and the ground-truth weight measured on a scale at the collection
point. The dataset targets vision-based weight estimation and bunch detection for smallholder
and estate harvest logistics.
Ringkasan: 31 tandan buah segar kelapa sawit varietas TANERA dari blok 303, difoto dari dua
sisi dan ditimbang langsung di lapangan. Setiap gambar disertai kotak… See the full description on the dataset page: https://huggingface.co/datasets/ULM-DS-Lab/Sawit-Weight.en-si-gemma3-translation-master-10ksawit-ttudataseten-si-translation-weblate-technical-1k
En Si Translation Weblate Technical 1K
Dataset Summary
English-Sinhala Technical and UI Localization dataset with ~1,000 rows targeting software interfaces, technical terminology, and static variables.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 1000
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and extracted from the… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-weblate-technical-1k.en-si-parallel-3k
Dataset Card for en-si-parallel-3k
Dataset Summary
The en-si-parallel-3k dataset is a high-quality, synthetically generated parallel corpus containing 3,000 English-Sinhala translation pairs. It is specifically designed for fine-tuning Large Language Models (LLMs) to enhance English-to-Sinhala translation capabilities and cross-lingual understanding.
Dataset Composition
The dataset is structured into 60 distinct batches of 50 examples each, covering… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-parallel-3k.
