CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NimanthaPerera /sinhala_dataset_vtextn<1K0 likes456 downloads14h agoHugging Face02Chamaka8 /serendip-cpt-sinhala Serendib LLM CPT Sinhala Corpus A large-scale, deduplicated, quality-filtered Sinhala plain-text corpus built for Continual Pre-Training (CPT) of large language models. This dataset was used to adapt Meta-LLaMA-3-8B to the Sinhala language domain as part of the Serendib LLM Honours Degree Research Project at the University of Central Lancashire (UCLan), 2025–2026. This is one of the largest openly published Sinhala NLP corpora available, containing 23,449,223 training documents… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/serendip-cpt-sinhala.texttext-generation10M<n<100M0 likes79 downloads6mo agoHugging Face03ihalage /sinhala-instruction-finetune-large Dataset Card for sinhala-instruction-finetune-large Sinhala instruction finetune (SIF) dataset contains high quality question-answer pairs in Sinhala language. It is an aggregate of translated English datasets using Google Translate API and several Sinhala datasets in the Hugging Face Datasets hub. SIF dataset has been compiled by transforming the datasets specified below into a common format. sinhala_eli5 sinhala-llm-dataset-llama-prompt-format alpaca-sinhala… See the full description on the dataset page: https://huggingface.co/datasets/ihalage/sinhala-instruction-finetune-large.textquestion-answering100K<n<1M2 likes51 downloads9mo agoHugging Face04kasunUdayanga /Sinhala_Annotation_Dataset Sinhala Named Entity Recognition (NER) Dataset - 85,000 Annotations Dataset Description This is a high-quality Named Entity Recognition (NER) dataset for the Sinhala language, consisting of approximately 85,000 annotations. The dataset was manually curated and annotated by a team of three students to support NLP research for low-resource languages. The data is sourced from diverse domains, including social media comments, news articles, and public domain texts, capturing… See the full description on the dataset page: https://huggingface.co/datasets/kasunUdayanga/Sinhala_Annotation_Dataset.texttoken-classification1K<n<10K4 likes40 downloads9mo agoHugging Face05Chamaka8 /Serendip-sft-sinhala Serendip-SFT-Sinhala Dataset 🇱🇰 📊 Dataset Summary Serendip-SFT-Sinhala is a large-scale Sinhala instruction-tuning dataset with 293,613 high-quality examples for supervised fine-tuning (SFT) of large language models. Created to train SerendipLLM, a Sinhala language model designed to excel at instruction-following, question-answering, summarization, and text classification. 🌟 Highlights 🇱🇰 293,613 Sinhala examples (largest Sinhala SFT dataset) 📚 4 task… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/Serendip-sft-sinhala.texttext-generation100K<n<1M0 likes34 downloads7mo agoHugging Face06sinhala-nlp /NSINAgated NSINa - A {N}ews Corpus for {Sin}hal{a} This repository introduces NSINA, a comprehensive news corpus of over 500,000 articles from popular Sinhala news websites. Alongside NSINA, with different subsets, we also introduce three Sinhala NLP tasks (1) News Media Identification (2) News Category Prediction and (3) News Headline Generation. The release of NSINA aims to provide a solution to challenges in adapting large language models to Sinhala, offering valuable benchmarks and… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/NSINA.text100K<n<1M4 likes28 downloads3y agoHugging Face07AyeshaKalpani98 /Questions_Answers_In_Sinhala_Language@misc{AyeshaKalpani_2024, title={Questions_Answers_In_Sinhala_Language}, author={Ayesha Kalpani}, year={2024}, url={}, } Questions_Answers_In_Sinhala_Language Dataset Description A dataset containing questions and answers in the Sinhala language. This dataset is intended for training and evaluating question-answering models in Sinhala. Dataset Details License This dataset is licensed under the MIT License. Task… See the full description on the dataset page: https://huggingface.co/datasets/AyeshaKalpani98/Questions_Answers_In_Sinhala_Language.textquestion-answeringn<1K0 likes27 downloads2y agoHugging Face08sh4lu-z /Sinhala-Mega-Corpus-v1 Sinhala Mega Corpus v1 Description English: Sinhala Mega Corpus v1 is a large-scale, high-quality merged dataset specifically designed for training Sinhala Large Language Models (LLMs) and Tokenizers. It combines several major open-source datasets into a single, unified format, providing a diverse range of linguistic patterns from web crawls, encyclopedic knowledge, and conversational data. සිංහල: Sinhala Mega Corpus v1 යනු සිංහල Large Language Models (LLM) සහ Tokenizers… See the full description on the dataset page: https://huggingface.co/datasets/sh4lu-z/Sinhala-Mega-Corpus-v1.texttext-generation100K<n<1M0 likes26 downloads7mo agoHugging Face09ihalage /sinhala-finetune-qa-eli5 Dataset Card for sinhala-finetune-qa-eli5 Sinhala question answering (QA) dataset contains a subset of the translated eli5 (explain like I'm 5) English dataset. eli5 is a crowdsourced dataset based mainly on the content from the subreddit r/explainlikeimfive. This is a forum where users post complex questions and other users provide simplified explanations. A subset of eli5 dataset (10k samples) has been machine translated to Sinhala language using the Google Cloud Translation API.… See the full description on the dataset page: https://huggingface.co/datasets/ihalage/sinhala-finetune-qa-eli5.textquestion-answering10K<n<100K2 likes23 downloads2y agoHugging Face10ov1n /science-sinhala-gce-olevel-2023-mcq Dataset Details This dataset contains 40 Science MCQ questions and answers in Sinhala language of the GCE Ordinary Level Science paper 2023. textquestion-answeringn<1K0 likes19 downloads2y agoHugging Face11sayururehan /sinhala-personas-lk-v0.2-gemini-1000-preview Sinhala-Personas-LK v0.2 Gemini 1000 Preview Sinhala-Personas-LK is a preview dataset of fully synthetic Sinhala persona records for Sri Lankan NLP research and evaluation. Version 0.1-preview Records 1000 synthetic records. Language Sinhala (si). Country context Sri Lanka (LK). Important limitations This preview version is generated from starter priors and LLM-generated text. It is not yet fully grounded… See the full description on the dataset page: https://huggingface.co/datasets/sayururehan/sinhala-personas-lk-v0.2-gemini-1000-preview.tabulartext-generation1K<n<10K0 likes17 downloads4mo agoHugging Face12ovinduG /gemma4-sinhala-cpt-evaltabularn<1K0 likes11 downloads3mo agoHugging Face13ov1n /sinhala-political-science-gce-alevel-2021-questionstextquestion-answeringn<1K0 likes9 downloads2y agoHugging Face14QuixiAI /SystemChat_SinhalaThis is a SystemChat dataset for Sinhala textn<1K6 likes8 downloads2y agoHugging Face15ov1n /sinhala-general-knowledgegated Dataset Details This dataset contains 220 general knowledge questions and answers in Sinhala language ona variety of domains. textquestion-answeringn<1K0 likes8 downloads2y agoHugging Face16ov1n /sinhala-geography-gce-alevel-2019textquestion-answeringn<1K0 likes8 downloads2y agoHugging Face17ov1n /sinhala-agriculture-gce-alevel-2021textquestion-answeringn<1K0 likes8 downloads2y agoHugging Face18ov1n /sinhala-geography-gce-alevel-2020textquestion-answeringn<1K0 likes6 downloads2y agoHugging Face19ov1n /sinhala-alevel-physics-questionsgated Dataset Details This dataset contains 20 physics questions and answers focused on Sinhala language. tabularquestion-answeringn<1K0 likes5 downloads2y agoHugging Face20pasindubg /sinhala-training-filestext10K<n<100K0 likes3 downloads11mo agoHugging Face21Isuru0x01 /sinhala_questions_answerstextn<1K0 likes2 downloads2y agoHugging Face22Sachin-Hansaka /SQAD-Sinhala_Question_Answering_DatasetgatedThis dataset is a back-translated version of the SQuAD 2.0 dataset, translated into Sinhala using the Google Cloud Translate API by Sachin Hansaka. Original dataset by the Stanford QA Group: https://rajpurkar.github.io/SQuAD-explorer/ Original work licensed under CC BY-SA 4.0. This Sinhala version © 2025 Sachin Hansaka, also licensed under CC BY-SA 4.0. 📚 Dataset Overview SQAD-Sinhala_Question_Answering_Dataset is a high-quality, back-translated version of the original SQuAD 2.0 dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sachin-Hansaka/SQAD-Sinhala_Question_Answering_Dataset.textquestion-answering100K<n<1M1 likes2 downloads1y agoHugging Face23Pamzyy /bbc_sinhalatext1K<n<10K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.