datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
awesome-dataset-sinhala
Mixed Sinhala Dataset (1M+ Rows) | මිශ්ර සිංහල දත්ත කට්ටලය
(Please find the English description below the Sinhala description)
🇬🇧 English
This is a comprehensive dataset containing over one million rows of Sinhala text data. It is highly suitable for training Artificial Intelligence (AI) models and conducting Natural Language Processing (NLP) research.
Dataset Details
Language: Sinhala (si)
Total Rows: 1,079,909
Format: Parquet (Optimized for Hugging… See the full description on the dataset page: https://huggingface.co/datasets/sh4lu-z/awesome-dataset-sinhala.UltraChat-Sinhala
Dataset Card for UltraChat-Sinhala
Dataset Description
UltraChat-Sinhala is a Sinhala (සිංහල) machine translation of
HuggingFaceH4/ultrachat_200k,
built to supervised-fine-tune Sinhala large language models. It preserves the
original dataset's structure, splits, and prompt_ids, so it is a drop-in
Sinhala counterpart to the English source.
The English dialogues were translated with
NLLB-200-3.3B
(eng_Latn → sin_Sinh) and then put through a Sinhala-specific cleaning… See the full description on the dataset page: https://huggingface.co/datasets/ThisenEkanayake/UltraChat-Sinhala.serendip-cpt-sinhala
Serendib LLM CPT Sinhala Corpus
A large-scale, deduplicated, quality-filtered Sinhala plain-text corpus built for
Continual Pre-Training (CPT) of large language models. This dataset was used to adapt
Meta-LLaMA-3-8B to the Sinhala language domain as part of the
Serendib LLM Honours Degree Research Project
at the University of Central Lancashire (UCLan), 2025–2026.
This is one of the largest openly published Sinhala NLP corpora available, containing
23,449,223 training documents… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/serendip-cpt-sinhala.Sinhala-Song-Lyrics
🎵 SynhalaAI — Ultimate Sinhala Song Lyrics Dataset (Gold Mix)
Dataset Description
The SynhalaAI Lyrics Corpus is a meticulously engineered, high-fidelity dataset of Sinhala song lyrics. It was designed specifically to train Large Language Models (LLMs) and advanced tokenizers on the poetic, colloquial, and structured linguistic patterns of the Sinhala language.
Unlike standard web-scraped datasets that are littered with English guitar chords, metadata, and HTML… See the full description on the dataset page: https://huggingface.co/datasets/SynhalaAI/Sinhala-Song-Lyrics.sinhala-corpus-c-diverse-1m
Diversity-Optimized Sinhala Corpus
A diversity-optimized subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model C) as part of a diversity-driven Sinhala language model adaptation study.
Corpus variants in this series:
Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline)
Minuri/sinhala-corpus-b-random-1m - Random subset (random baseline)… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-c-diverse-1m.sinhala-corpus-wikipedia
Sinhala Raw Sentences - Wikipedia
Raw Sinhala sentences extracted and sentence-split from the wikimedia/wikipedia dataset (Sinhala subset). This is an intermediate dataset used in the construction of Minuri/diverse_sinhala_dataset.
Dataset Structure
Column
Description
text
Raw Sinhala sentence
source
Source identifier (wikipedia)
Split
Rows
train
517,246
Pipeline Position
wikimedia/wikipedia → this repo →… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-wikipedia.sinhala-instruction-finetune-large
Dataset Card for sinhala-instruction-finetune-large
Sinhala instruction finetune (SIF) dataset contains high quality question-answer pairs in Sinhala language. It is an aggregate of translated English datasets using Google Translate API and several Sinhala datasets in the
Hugging Face Datasets hub. SIF dataset has been compiled by transforming the datasets specified below into a common format.
sinhala_eli5
sinhala-llm-dataset-llama-prompt-format
alpaca-sinhala… See the full description on the dataset page: https://huggingface.co/datasets/ihalage/sinhala-instruction-finetune-large.sinhala-english-singlish-translation
Sinhala–English–Singlish Translation Dataset
A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations.
📋 Table of Contents
Dataset Overview
Installation
Quick Start
Dataset Structure
Usage Examples
Citation
License
Credits
Dataset Overview
Description: 34,500 aligned triplets of
Sinhala (native script)
English (human translation)
Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sinhala-english-singlish-translation.SinhalaDentalQnAsinhala-sft-dataset
Sinhala Supervised Fine-Tuning Dataset
A merged Sinhala instruction-following dataset of 213,703 pairs, used for Supervised Fine-Tuning (SFT) of continually pretrained LLaMA 3.2 1B variants. Constructed as part of a diversity-driven Sinhala language model adaptation study.
Dataset Description
This dataset merges three existing Sinhala instruction datasets into a unified resource for SFT. It follows the standard Alpaca-style instruction–input–output format and covers a… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-sft-dataset.sinhala-text-dataset
Sinhala Continuous Pretraining Corpus
Curated by HelaAI
Dataset Summary
This dataset is a Sinhala-language text corpus assembled for continuous pretraining of language models. It combines multiple sources into a single, cleaned, block-structured corpus:
News articles — Sinhala news text extracted from the article_sinhala field of Hamza-Ziyard/CNN-Daily-Mail-Sinhala.
O/L Sinhala Buddhism — Sinhala-medium educational text covering the GCE Ordinary Level (O/L)… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/sinhala-text-dataset.sinhala-detoxification-dimuthu-subset
📊 Dataset Card for Sinhala Detoxification - Dimuthu's Subset
📝 Dataset Description
Repository: dimuthulk/sinhala-detoxification-dimuthu-subset
Language(s) (NLP): Sinhala (si)
License: Apache 2.0
📋 Dataset Summary
This dataset represents the individual data collection and curation contribution of Dimuthu Rathnayaka for the overarching research project, "Sinhala Offensive Text Detoxification Pipeline".
It was developed as part of the academic… See the full description on the dataset page: https://huggingface.co/datasets/dimuthulk/sinhala-detoxification-dimuthu-subset.sinhala-text-detoxification
📊 Dataset Card for Sinhala Text Detoxification Dataset
📝 Dataset Description
Repository: dimuthulk/sinhala-text-detoxification
Language(s) (NLP): Sinhala (si)
License: Apache 2.0
📋 Dataset Summary
This dataset serves as the primary generative dataset for the second phase of the "Sinhala Offensive Text Detoxification Pipeline" research project conducted at the University of Kelaniya (Electronics and computer science degree program).
It contains… See the full description on the dataset page: https://huggingface.co/datasets/dimuthulk/sinhala-text-detoxification.diverse_sinhala_dataset
Diverse Sinhala Dataset
A large-scale, cleaned, deduplicated, and domain-classified Sinhala text corpus compiled for continual pretraining of large language models. Constructed as part of a diversity-driven Sinhala language model adaptation study.
This repository serves as pipeline storage for the full corpus construction process, from merged cleaned sentences through to the final high-confidence domain-classified corpus.
Files
File
Rows
Columns
Description… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/diverse_sinhala_dataset.sinhala-corpus-madlad400
Sinhala Raw Sentences - MADLAD-400
Raw Sinhala sentences extracted and sentence-split from the allenai/MADLAD-400 dataset. This is an intermediate dataset used in the construction of Minuri/diverse_sinhala_dataset.
Dataset Structure
Column
Description
text
Raw Sinhala sentence
source
Source identifier (madlad)
Split
Rows
train
7,281,026
Pipeline Position
allenai/MADLAD-400 → this repo → Minuri/madlad_cleaned_version →… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-madlad400.sinhala-corpus-culturax
Sinhala Raw Sentences - CulturaX
Raw Sinhala sentences extracted and sentence-split from the uonlp/CulturaX dataset. This is an intermediate dataset used in the construction of Minuri/diverse_sinhala_dataset.
Dataset Structure
Column
Description
text
Raw Sinhala sentence
source
Source identifier (culturax)
Split
Rows
train
4,707,451
Pipeline Position
uonlp/CulturaX → this repo → Minuri/culturax_cleaned_version →… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-culturax.sinhala-test-set-50k
Sinhala Test Set - 50K Sentences
A held-out Sinhala test set of 50,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for perplexity evaluation of three continually pretrained LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study.
Dataset Description
This test set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased perplexity measurement. It covers multiple… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-test-set-50k.openslr-sinhala-synthetic-spell-errors-quarter
Sinhala Dyslexic Spelling Correction Dataset
Dataset Description
This dataset contains Sinhala and code-mixed (Sinhala-English) text pairs for training spelling correction models, specifically designed to address dyslexia-like spelling errors.
Features
dyslexic_sentence: Input text with dyslexia-like spelling errors (string)
correct_sentence: Corrected output text (string)
Dataset Statistics
Split
Samples
Train
37,056
Test
9,265… See the full description on the dataset page: https://huggingface.co/datasets/SPEAK-PP/openslr-sinhala-synthetic-spell-errors-quarter.sinhala-corpus-b-random-1m
Randomly Curated Sinhala Corpus
A randomly sampled subset of 1M Sinhala sentences from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model B) as part of a diversity-driven Sinhala language model adaptation study.
Corpus variants in this series:
Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline)
Minuri/sinhala-corpus-b-random-1m - Random subset (random baseline) - this repo
Minuri/sinhala-corpus-c-diverse-1m… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-b-random-1m.cnn-dailymail-sinhala-continuous-pretrain
CNN DailyMail Sinhala Continuous Pretraining Dataset
Dataset Description
This dataset is designed for continuous pretraining of Sinhala Small Language Models (SLMs) and Large Language Models (LLMs).
The dataset was created by processing the original Sinhala news articles from:
CNN Daily Mail Sinhala Dataset
The article_sinhala field from the original dataset was extracted, cleaned, and concatenated into larger continuous text blocks suitable for language model… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/cnn-dailymail-sinhala-continuous-pretrain.SinhalaCorpusLargeSerendip-sft-sinhala
Serendip-SFT-Sinhala Dataset 🇱🇰
📊 Dataset Summary
Serendip-SFT-Sinhala is a large-scale Sinhala instruction-tuning dataset with 293,613 high-quality examples for supervised fine-tuning (SFT) of large language models.
Created to train SerendipLLM, a Sinhala language model designed to excel at instruction-following, question-answering, summarization, and text classification.
🌟 Highlights
🇱🇰 293,613 Sinhala examples (largest Sinhala SFT dataset)
📚 4 task… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/Serendip-sft-sinhala.SinhalaWikipediaArticlesQuestions_Answers_In_Sinhala_Language@misc{AyeshaKalpani_2024,
title={Questions_Answers_In_Sinhala_Language},
author={Ayesha Kalpani},
year={2024},
url={},
}
Questions_Answers_In_Sinhala_Language
Dataset Description
A dataset containing questions and answers in the Sinhala language. This dataset is intended for training and evaluating question-answering models in Sinhala.
Dataset Details
License
This dataset is licensed under the MIT License.
Task… See the full description on the dataset page: https://huggingface.co/datasets/AyeshaKalpani98/Questions_Answers_In_Sinhala_Language.sinhala-corpus-a-news-1m
News-Only Sinhala Corpus
A news-domain subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model A) as part of a diversity-driven Sinhala language model adaptation study at the Informatics Institute of Technology (IIT), Colombo, affiliated with Robert Gordon University (RGU).
Corpus variants in this series:
Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline) - this repo… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-a-news-1m.SinhalaWikipediaArticlesSinhala-Mega-Corpus-v1
Sinhala Mega Corpus v1
Description
English:
Sinhala Mega Corpus v1 is a large-scale, high-quality merged dataset specifically designed for training Sinhala Large Language Models (LLMs) and Tokenizers. It combines several major open-source datasets into a single, unified format, providing a diverse range of linguistic patterns from web crawls, encyclopedic knowledge, and conversational data.
සිංහල:
Sinhala Mega Corpus v1 යනු සිංහල Large Language Models (LLM) සහ Tokenizers… See the full description on the dataset page: https://huggingface.co/datasets/sh4lu-z/Sinhala-Mega-Corpus-v1.Sinhala-News-Wiki-text-corpus
Sinhala-News-Wiki-Text-Corpus
Containing news articles from various Sinhala news sites along with Sinhala Wikipedia pages.
Dataset Overview
Language: Sinhala (සිංහල)
Content: Sinhala news articles from various sites
Data format: Parquet
Number of Records: 18,201 rows (as per current size)
Dataset Structure
Each record consists of the following fields:
category: The news category (e.g., "Other-news, Local-news, wiki, International-news").
site: The site's… See the full description on the dataset page: https://huggingface.co/datasets/janani-rane/Sinhala-News-Wiki-text-corpus.sinhala-articles
Sinhala Articles Dataset
A large-scale, high-quality Sinhala text corpus curated from diverse sources including news articles, Wikipedia entries, and general web content. This dataset is designed to support a wide range of Sinhala Natural Language Processing (NLP) tasks.
📊 Dataset Overview
Name: Navanjana/sinhala-articles
Total Samples: 2,148,688
Languages: Sinhala (si)
Features:
text: A single column containing Sinhala text passages.
Size: Approximately 1M < n <… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/sinhala-articles.sinhala-finetune-qa-eli5
Dataset Card for sinhala-finetune-qa-eli5
Sinhala question answering (QA) dataset contains a subset of the translated eli5 (explain like I'm 5) English dataset. eli5 is a crowdsourced dataset based mainly on the content from the subreddit r/explainlikeimfive.
This is a forum where users post complex questions and other users provide simplified explanations.
A subset of eli5 dataset (10k samples) has been machine translated to Sinhala language using the Google Cloud Translation API.… See the full description on the dataset page: https://huggingface.co/datasets/ihalage/sinhala-finetune-qa-eli5.
