datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sinhala-flansinhala_dataset_vsinhala-22gb-cleaned-datasetsinhala_dataset_sanitizedmini-sinhala-flanSinhalaCXSOLD
SOLD - A Benchmark for Sinhala Offensive Language Identification
In this repository, we introduce the {S}inhala {O}ffensive {L}anguage {D}ataset (SOLD) and present multiple experiments on this dataset. SOLD is a manually annotated dataset containing 10,000 posts from Twitter annotated as offensive and not offensive at both sentence-level and token-level. SOLD is the largest offensive language dataset compiled for Sinhala. We also introduce SemiSOLD, a larger dataset containing more… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/SOLD.sinhala_storiesopenslr-sinhala-asrsinhala-openslr-111hsinhala_sentences_rawopenslr-sinhala-asr-normlarge-sinhala-asr-datasetsinhala_30m
Dataset Card for "sinhala_30m"
More Information needed
openslr-sinhala-asr-depricated-versionsinhala-tts-dataset-archive-20260429-082457
Sinhala TTS Dataset
Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare.
Stats
Metric
Value
Utterances
218
Train
208
Val
10
Hours
0.51
Mean duration
8.5s
Sample rate
22050 Hz
Pipeline
Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 ->
Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB)
Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.sinhala-ocr-lk-acts-1010
🇱🇰 Sinhala OCR - Sri Lankan Acts Dataset
Dataset Description
This dataset contains 1,010 scanned document images of Sri Lankan legal acts (1980s-2010s) in Sinhala language with ground truth text annotations for Optical Character Recognition (OCR) training and evaluation.
Key Features
✅ High-quality scanned document images
✅ Professionally corrected ground truth text
✅ Year-wise metadata for temporal analysis
✅ Pre-split into train/eval/test sets… See the full description on the dataset page: https://huggingface.co/datasets/avishadilhara/sinhala-ocr-lk-acts-1010.SinhalaNewsClassification
SinhalaNewsClassification
An MTEB dataset
Massive Text Embedding Benchmark
This file contains news texts (sentences) belonging to 5 different news categories (political, business, technology, sports and Entertainment). The original dataset was released by Nisansa de Silva (Sinhala Text Classification: Observations from the Perspective of a Resource Poor Language, 2015).
Task category
t2c
Domains
News, Written
Reference… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SinhalaNewsClassification.SinhalaASR-testawesome-dataset-sinhala
Mixed Sinhala Dataset (1M+ Rows) | මිශ්ර සිංහල දත්ත කට්ටලය
(Please find the English description below the Sinhala description)
🇬🇧 English
This is a comprehensive dataset containing over one million rows of Sinhala text data. It is highly suitable for training Artificial Intelligence (AI) models and conducting Natural Language Processing (NLP) research.
Dataset Details
Language: Sinhala (si)
Total Rows: 1,079,909
Format: Parquet (Optimized for Hugging… See the full description on the dataset page: https://huggingface.co/datasets/sh4lu-z/awesome-dataset-sinhala.SinhalaASR-1000sinhala_synthetic_ocr_news_largeCreated from https://www.kaggle.com/code/ransakaravihara/sinhala-ocr-image-creation.
To cite the dataset
@misc{ransaka_r._2026,
author = { Ransaka R. },
title = { sinhala_synthetic_ocr_news_large (Revision bc52307) },
year = 2026,
url = { https://huggingface.co/datasets/edifier99/sinhala_synthetic_ocr_news_large },
doi = { 10.57967/hf/9748 },
publisher = { Hugging Face }
}
sinhala_dataset_59m
Dataset Card for "sinhala_dataset_59m"
More Information needed
sentencified_v1_sinhala_30m
Dataset Card for "sentencified_v1_sinhala_30m"
More Information needed
serendip-cpt-sinhala
Serendib LLM CPT Sinhala Corpus
A large-scale, deduplicated, quality-filtered Sinhala plain-text corpus built for
Continual Pre-Training (CPT) of large language models. This dataset was used to adapt
Meta-LLaMA-3-8B to the Sinhala language domain as part of the
Serendib LLM Honours Degree Research Project
at the University of Central Lancashire (UCLan), 2025–2026.
This is one of the largest openly published Sinhala NLP corpora available, containing
23,449,223 training documents… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/serendip-cpt-sinhala.named-entity-recognitionNSINA-Categories
Sinhala News Category Prediction
This is a text classification task created with the NSINA dataset. This dataset is also released with the same license as NSINA.
Data
Data can be loaded into pandas dataframes using the following code.
from datasets import Dataset
from datasets import load_dataset
train = Dataset.to_pandas(load_dataset('sinhala-nlp/NSINA-Categories', split='train'))
test = Dataset.to_pandas(load_dataset('sinhala-nlp/NSINA-Categories', split='test'))… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/NSINA-Categories.Sinhala-News-Category-classificationThis file contains news texts (sentences) belonging to 5 different news categories (political, business, technology, sports and Entertainment). The original dataset was released by Nisansa de Silva (Sinhala Text Classification: Observations from the Perspective of a Resource Poor Language, 2015). The original dataset is processed and cleaned of single word texts, English only sentences etc.
If you use this dataset, please cite {Nisansa de Silva, Sinhala Text Classification: Observations from… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Category-classification.Augmented_SinhalatoRomanizedSinhala_Dataset
Sinhala Romanized Dataset
This dataset contains Sinhala text with romanized (transliterated) versions, created using publicly available Sinhala data sources.
Dataset Description
The Augmented Sinhala to Romanized Sinhala Dataset provides paired examples of Sinhala text and their corresponding romanized transliterations. This dataset aims to facilitate research in Sinhala language processing, particularly for applications that require romanized representations of Sinhala… See the full description on the dataset page: https://huggingface.co/datasets/deshanksuman/Augmented_SinhalatoRomanizedSinhala_Dataset.SiMTEB-NHP
