datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sinhala-english-singlish-translation
Sinhala–English–Singlish Translation Dataset
A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations.
📋 Table of Contents
Dataset Overview
Installation
Quick Start
Dataset Structure
Usage Examples
Citation
License
Credits
Dataset Overview
Description: 34,500 aligned triplets of
Sinhala (native script)
English (human translation)
Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sinhala-english-singlish-translation.sinhala-articles
Sinhala Articles Dataset
A large-scale, high-quality Sinhala text corpus curated from diverse sources including news articles, Wikipedia entries, and general web content. This dataset is designed to support a wide range of Sinhala Natural Language Processing (NLP) tasks.
📊 Dataset Overview
Name: Navanjana/sinhala-articles
Total Samples: 2,148,688
Languages: Sinhala (si)
Features:
text: A single column containing Sinhala text passages.
Size: Approximately 1M < n <… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/sinhala-articles.sinhala-spell-correction-dataset
Sinhala Spell Correction Dataset
A Sinhala spell correction dataset used for training and evaluating neural spell correction models as part of the LMSpell project.
Dataset Description
This dataset combines data from previously published Sinhala spell correction resources and applies additional cleaning to improve its suitability for training neural spell correction models.
The dataset originates from the benchmark introduced by Sonnadara et al. (2021) and was… See the full description on the dataset page: https://huggingface.co/datasets/lm-spell/sinhala-spell-correction-dataset.Sinhala-Wikianonymized-sinhala-letter-corpus
Anonymized Sinhala Official Letter Corpus
A small, hand-curated corpus of 151 formal Sinhala letters, fully anonymized with
bracketed placeholders. It is intended for training and evaluating models that generate,
complete, or classify Sinhala official correspondence — a task with very little public
training data.
Dataset at a glance
Examples
151
Language
Sinhala (si)
Register
Formal throughout
Letter length
42–240 words (median 108, mean 114)… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/anonymized-sinhala-letter-corpus.NSINA-Headlines
Sinhala Headline Generation
This is a text generation task created with the NSINA dataset. This dataset is also released with the same license as NSINA. The objective of the task is to generate news headlines based on the provided news content.
Data
We used the same instances from NSINA 1.0 as all the news articles had headlines. We divided this dataset into a training and test set following a 0.8 split.
Data can be loaded into pandas dataframes using the following code.… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/NSINA-Headlines.
