datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mini-sinhala-flansinhala-corpus-c-diverse-1m
Diversity-Optimized Sinhala Corpus
A diversity-optimized subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model C) as part of a diversity-driven Sinhala language model adaptation study.
Corpus variants in this series:
Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline)
Minuri/sinhala-corpus-b-random-1m - Random subset (random baseline)… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-c-diverse-1m.SemiSOLD
SOLD - A Benchmark for Sinhala Offensive Language Identification
In this repository, we introduce the {S}inhala {O}ffensive {L}anguage {D}ataset (SOLD) and present multiple experiments on this dataset. SOLD is a manually annotated dataset containing 10,000 posts from Twitter annotated as offensive and not offensive at both sentence-level and token-level. SOLD is the largest offensive language dataset compiled for Sinhala. We also introduce SemiSOLD, a larger dataset containing more… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/SemiSOLD.sinhala-vqa-dataset
Sinhala VQA Dataset
A Sinhala-language Visual Question Answering dataset of 37,318 QA pairs, constructed by translating Visual Genome QA annotations into Sinhala using gemini-3-flash-preview. This dataset was developed as part of research on benchmarking and adapting compact multimodal models for Sinhala VQA under low-resource conditions.
Dataset Summary
Split
Samples
Train
33,409
Validation
2,909
Test
1,000
Total
37,318
Schema
Each row… See the full description on the dataset page: https://huggingface.co/datasets/Siluni/sinhala-vqa-dataset.Sinhala-News-Source-classificationThis dataset contains Sinhala news headlines extracted from 9 news sources (websites) (Sri Lanka Army, Dinamina, GossipLanka, Hiru, ITN, Lankapuwath, NewsLK,
Newsfirst, World Socialist Web Site-Sinhala). This is a processed version of the corpus created by Sachintha, D., Piyarathna, L., Rajitha, C., and Ranathunga, S. (2021). Exploiting parallel corpora to improve multilingual embedding based document and sentence alignment. Single word sentences, invalid characters have been removed from the… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Source-classification.sinhala-test-set-50k
Sinhala Test Set - 50K Sentences
A held-out Sinhala test set of 50,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for perplexity evaluation of three continually pretrained LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study.
Dataset Description
This test set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased perplexity measurement. It covers multiple… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-test-set-50k.sinhala-summarization-dataset
Sinhala Text Summarization Dataset
Dataset Description
This dataset is a Sinhala text summarization dataset created for research in low-resource language summarization. The dataset contains 2,493 Sinhala article-summary pairs collected from diverse publicly accessible Sinhala online sources.
This repository contains a Sinhala article-summary dataset introduced in the following IEEE conference publication:
Sinhala Automatic Text Summarization: Dataset Creation and… See the full description on the dataset page: https://huggingface.co/datasets/hans1k/sinhala-summarization-dataset.sinhala-corpus-b-random-1m
Randomly Curated Sinhala Corpus
A randomly sampled subset of 1M Sinhala sentences from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model B) as part of a diversity-driven Sinhala language model adaptation study.
Corpus variants in this series:
Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline)
Minuri/sinhala-corpus-b-random-1m - Random subset (random baseline) - this repo
Minuri/sinhala-corpus-c-diverse-1m… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-b-random-1m.Sinhala-News-Wiki-text-corpus
Sinhala-News-Wiki-Text-Corpus
Containing news articles from various Sinhala news sites along with Sinhala Wikipedia pages.
Dataset Overview
Language: Sinhala (සිංහල)
Content: Sinhala news articles from various sites
Data format: Parquet
Number of Records: 18,201 rows (as per current size)
Dataset Structure
Each record consists of the following fields:
category: The news category (e.g., "Other-news, Local-news, wiki, International-news").
site: The site's… See the full description on the dataset page: https://huggingface.co/datasets/janani-rane/Sinhala-News-Wiki-text-corpus.sinhala-corpus-a-news-1m
News-Only Sinhala Corpus
A news-domain subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model A) as part of a diversity-driven Sinhala language model adaptation study at the Informatics Institute of Technology (IIT), Colombo, affiliated with Robert Gordon University (RGU).
Corpus variants in this series:
Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline) - this repo… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-a-news-1m.SinhalaMMLU
SinhalaMMLU
We introduce SinhalaMMLU, the first multiple-choice question answering benchmark designed specifically for Sinhala, a low-resource language.The dataset contains over 7,000 questions spanning secondary to collegiate education levels, aligned with the Sri Lankan national curriculum.It covers six domains and 30 subjects, encompassing both general academic topics and culturally grounded knowledge. We evaluated
26 large language models (LLMs) on SinhalaMMLU and observed that… See the full description on the dataset page: https://huggingface.co/datasets/naist-nlp/SinhalaMMLU.sinhala-personas-lk-v0.2-gemini-1000-preview
Sinhala-Personas-LK v0.2 Gemini 1000 Preview
Sinhala-Personas-LK is a preview dataset of fully synthetic Sinhala persona records for Sri Lankan NLP research and evaluation.
Version
0.1-preview
Records
1000 synthetic records.
Language
Sinhala (si).
Country context
Sri Lanka (LK).
Important limitations
This preview version is generated from starter priors and LLM-generated text. It is not yet fully grounded… See the full description on the dataset page: https://huggingface.co/datasets/sayururehan/sinhala-personas-lk-v0.2-gemini-1000-preview.Testing_audio_text_pairs_Sinhala_v1sinhala-news-sentiment-classification
Dataset Card for "sinhala-news-sentiment-classification"
More Information needed
sinhala-poems-v1
Sinhala Poems (Filtered)
Curated Sinhala poem blocks extracted from web blogs using a verse-shaped heuristic (v3.4) with weak attributes (theme/mood/style) and stats.
Columns
text: full original block (cleaned)
snippet: first stanza or 8 lines
url: source URL
subtype: poem | song_like | promo_like | unknown
keep: boolean accepted by filter
theme, mood, style, length_class: weak labels
ps_*: structure stats (floats)
Filtering summary
Boilerplate/HTML… See the full description on the dataset page: https://huggingface.co/datasets/manthilaffs/sinhala-poems-v1.facebook-sinhala-unicode-hate-speechgemma4-sinhala-cpt-evalakura-sinhala-dyslexic-writing-patterns
Akura Sinhala Dyslexic Writing Patterns Dataset
Overview
This dataset provides a sentence-level, feature-augmented corpus for diagnosing dyslexic writing patterns in Sinhala.Each instance consists of a dyslexic sentence, its corresponding clean reference sentence, a set of explicit character-level error features, and a dominant dyslexic writing pattern label.
Unlike correction-focused datasets, this corpus is designed for diagnostic classification, enabling models to… See the full description on the dataset page: https://huggingface.co/datasets/akura-official/akura-sinhala-dyslexic-writing-patterns.sinhala-validation-set-10k
Sinhala Validation Set - 10K Sentences
A held-out Sinhala validation set of 10,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for monitoring validation loss during continual pretraining of three LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study.
Dataset Description
This validation set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased validation loss… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-validation-set-10k.SinhalaMMLU-Complete
SinhalaMMLU
We introduce SinhalaMMLU, the first multiple-choice question answering benchmark designed specifically for Sinhala, a low-resource language.The dataset contains over 7,000 questions spanning secondary to collegiate education levels, aligned with the Sri Lankan national curriculum.It covers six domains and 30 subjects, encompassing both general academic topics and culturally grounded knowledge. We evaluated
26 large language models (LLMs) on SinhalaMMLU and observed… See the full description on the dataset page: https://huggingface.co/datasets/naist-nlp/SinhalaMMLU-Complete.FacebookDecadeCorporasinhala-alevel-physics-questions
Dataset Details
This dataset contains 20 physics questions and answers focused on Sinhala language.
sinhala_YouTube_Emotion_Splitssinhala-news-sentiment-classification
Dataset Card for "sinhala-news-sentiment-classification"
More Information needed
sinhala_YouTube_Comment_Processed
