datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sinhala_dataset_vserendip-cpt-sinhala
Serendib LLM CPT Sinhala Corpus
A large-scale, deduplicated, quality-filtered Sinhala plain-text corpus built for
Continual Pre-Training (CPT) of large language models. This dataset was used to adapt
Meta-LLaMA-3-8B to the Sinhala language domain as part of the
Serendib LLM Honours Degree Research Project
at the University of Central Lancashire (UCLan), 2025–2026.
This is one of the largest openly published Sinhala NLP corpora available, containing
23,449,223 training documents… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/serendip-cpt-sinhala.sinhala-instruction-finetune-large
Dataset Card for sinhala-instruction-finetune-large
Sinhala instruction finetune (SIF) dataset contains high quality question-answer pairs in Sinhala language. It is an aggregate of translated English datasets using Google Translate API and several Sinhala datasets in the
Hugging Face Datasets hub. SIF dataset has been compiled by transforming the datasets specified below into a common format.
sinhala_eli5
sinhala-llm-dataset-llama-prompt-format
alpaca-sinhala… See the full description on the dataset page: https://huggingface.co/datasets/ihalage/sinhala-instruction-finetune-large.Sinhala_Annotation_Dataset
Sinhala Named Entity Recognition (NER) Dataset - 85,000 Annotations
Dataset Description
This is a high-quality Named Entity Recognition (NER) dataset for the Sinhala language, consisting of approximately 85,000 annotations. The dataset was manually curated and annotated by a team of three students to support NLP research for low-resource languages.
The data is sourced from diverse domains, including social media comments, news articles, and public domain texts, capturing… See the full description on the dataset page: https://huggingface.co/datasets/kasunUdayanga/Sinhala_Annotation_Dataset.Serendip-sft-sinhala
Serendip-SFT-Sinhala Dataset 🇱🇰
📊 Dataset Summary
Serendip-SFT-Sinhala is a large-scale Sinhala instruction-tuning dataset with 293,613 high-quality examples for supervised fine-tuning (SFT) of large language models.
Created to train SerendipLLM, a Sinhala language model designed to excel at instruction-following, question-answering, summarization, and text classification.
🌟 Highlights
🇱🇰 293,613 Sinhala examples (largest Sinhala SFT dataset)
📚 4 task… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/Serendip-sft-sinhala.NSINA
NSINa - A {N}ews Corpus for {Sin}hal{a}
This repository introduces NSINA, a comprehensive news corpus of over 500,000 articles from popular Sinhala news websites. Alongside NSINA, with different subsets, we also introduce three Sinhala NLP tasks (1) News Media Identification (2) News Category Prediction and (3) News Headline Generation. The release of NSINA aims to provide a solution to challenges in adapting large language models to Sinhala, offering valuable benchmarks and… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/NSINA.Questions_Answers_In_Sinhala_Language@misc{AyeshaKalpani_2024,
title={Questions_Answers_In_Sinhala_Language},
author={Ayesha Kalpani},
year={2024},
url={},
}
Questions_Answers_In_Sinhala_Language
Dataset Description
A dataset containing questions and answers in the Sinhala language. This dataset is intended for training and evaluating question-answering models in Sinhala.
Dataset Details
License
This dataset is licensed under the MIT License.
Task… See the full description on the dataset page: https://huggingface.co/datasets/AyeshaKalpani98/Questions_Answers_In_Sinhala_Language.Sinhala-Mega-Corpus-v1
Sinhala Mega Corpus v1
Description
English:
Sinhala Mega Corpus v1 is a large-scale, high-quality merged dataset specifically designed for training Sinhala Large Language Models (LLMs) and Tokenizers. It combines several major open-source datasets into a single, unified format, providing a diverse range of linguistic patterns from web crawls, encyclopedic knowledge, and conversational data.
සිංහල:
Sinhala Mega Corpus v1 යනු සිංහල Large Language Models (LLM) සහ Tokenizers… See the full description on the dataset page: https://huggingface.co/datasets/sh4lu-z/Sinhala-Mega-Corpus-v1.sinhala-finetune-qa-eli5
Dataset Card for sinhala-finetune-qa-eli5
Sinhala question answering (QA) dataset contains a subset of the translated eli5 (explain like I'm 5) English dataset. eli5 is a crowdsourced dataset based mainly on the content from the subreddit r/explainlikeimfive.
This is a forum where users post complex questions and other users provide simplified explanations.
A subset of eli5 dataset (10k samples) has been machine translated to Sinhala language using the Google Cloud Translation API.… See the full description on the dataset page: https://huggingface.co/datasets/ihalage/sinhala-finetune-qa-eli5.science-sinhala-gce-olevel-2023-mcq
Dataset Details
This dataset contains 40 Science MCQ questions and answers in Sinhala language of the GCE Ordinary Level Science paper 2023.
sinhala-personas-lk-v0.2-gemini-1000-preview
Sinhala-Personas-LK v0.2 Gemini 1000 Preview
Sinhala-Personas-LK is a preview dataset of fully synthetic Sinhala persona records for Sri Lankan NLP research and evaluation.
Version
0.1-preview
Records
1000 synthetic records.
Language
Sinhala (si).
Country context
Sri Lanka (LK).
Important limitations
This preview version is generated from starter priors and LLM-generated text. It is not yet fully grounded… See the full description on the dataset page: https://huggingface.co/datasets/sayururehan/sinhala-personas-lk-v0.2-gemini-1000-preview.gemma4-sinhala-cpt-evalsinhala-political-science-gce-alevel-2021-questionsSystemChat_SinhalaThis is a SystemChat dataset for Sinhala
sinhala-general-knowledge
Dataset Details
This dataset contains 220 general knowledge questions and answers in Sinhala language ona variety of domains.
sinhala-geography-gce-alevel-2019sinhala-agriculture-gce-alevel-2021sinhala-geography-gce-alevel-2020sinhala-alevel-physics-questions
Dataset Details
This dataset contains 20 physics questions and answers focused on Sinhala language.
sinhala-training-filessinhala_questions_answersSQAD-Sinhala_Question_Answering_DatasetThis dataset is a back-translated version of the SQuAD 2.0 dataset, translated into Sinhala using the Google Cloud Translate API by Sachin Hansaka.
Original dataset by the Stanford QA Group: https://rajpurkar.github.io/SQuAD-explorer/
Original work licensed under CC BY-SA 4.0.
This Sinhala version © 2025 Sachin Hansaka, also licensed under CC BY-SA 4.0.
📚 Dataset Overview
SQAD-Sinhala_Question_Answering_Dataset is a high-quality, back-translated version of the original SQuAD 2.0 dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sachin-Hansaka/SQAD-Sinhala_Question_Answering_Dataset.bbc_sinhala
