CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sh4lu-z /awesome-dataset-sinhala Mixed Sinhala Dataset (1M+ Rows) | මිශ්‍ර සිංහල දත්ත කට්ටලය (Please find the English description below the Sinhala description) 🇬🇧 English This is a comprehensive dataset containing over one million rows of Sinhala text data. It is highly suitable for training Artificial Intelligence (AI) models and conducting Natural Language Processing (NLP) research. Dataset Details Language: Sinhala (si) Total Rows: 1,079,909 Format: Parquet (Optimized for Hugging… See the full description on the dataset page: https://huggingface.co/datasets/sh4lu-z/awesome-dataset-sinhala.texttext-generation1M<n<10M1 likes95 downloads7mo agoHugging Face02ThisenEkanayake /UltraChat-Sinhala Dataset Card for UltraChat-Sinhala Dataset Description UltraChat-Sinhala is a Sinhala (සිංහල) machine translation of HuggingFaceH4/ultrachat_200k, built to supervised-fine-tune Sinhala large language models. It preserves the original dataset's structure, splits, and prompt_ids, so it is a drop-in Sinhala counterpart to the English source. The English dialogues were translated with NLLB-200-3.3B (eng_Latn → sin_Sinh) and then put through a Sinhala-specific cleaning… See the full description on the dataset page: https://huggingface.co/datasets/ThisenEkanayake/UltraChat-Sinhala.text-generation100K<n<1M0 likes79 downloads3mo agoHugging Face03Chamaka8 /serendip-cpt-sinhala Serendib LLM CPT Sinhala Corpus A large-scale, deduplicated, quality-filtered Sinhala plain-text corpus built for Continual Pre-Training (CPT) of large language models. This dataset was used to adapt Meta-LLaMA-3-8B to the Sinhala language domain as part of the Serendib LLM Honours Degree Research Project at the University of Central Lancashire (UCLan), 2025–2026. This is one of the largest openly published Sinhala NLP corpora available, containing 23,449,223 training documents… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/serendip-cpt-sinhala.texttext-generation10M<n<100M0 likes69 downloads6mo agoHugging Face04SynhalaAI /Sinhala-Song-Lyrics 🎵 SynhalaAI — Ultimate Sinhala Song Lyrics Dataset (Gold Mix) Dataset Description The SynhalaAI Lyrics Corpus is a meticulously engineered, high-fidelity dataset of Sinhala song lyrics. It was designed specifically to train Large Language Models (LLMs) and advanced tokenizers on the poetic, colloquial, and structured linguistic patterns of the Sinhala language. Unlike standard web-scraped datasets that are littered with English guitar chords, metadata, and HTML… See the full description on the dataset page: https://huggingface.co/datasets/SynhalaAI/Sinhala-Song-Lyrics.texttext-generation1K<n<10K2 likes54 downloads3mo agoHugging Face05Minuri /sinhala-corpus-c-diverse-1m Diversity-Optimized Sinhala Corpus A diversity-optimized subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model C) as part of a diversity-driven Sinhala language model adaptation study. Corpus variants in this series: Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline) Minuri/sinhala-corpus-b-random-1m - Random subset (random baseline)… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-c-diverse-1m.tabulartext-generation1M<n<10M0 likes53 downloads6mo agoHugging Face06Minuri /sinhala-corpus-wikipedia Sinhala Raw Sentences - Wikipedia Raw Sinhala sentences extracted and sentence-split from the wikimedia/wikipedia dataset (Sinhala subset). This is an intermediate dataset used in the construction of Minuri/diverse_sinhala_dataset. Dataset Structure Column Description text Raw Sinhala sentence source Source identifier (wikipedia) Split Rows train 517,246 Pipeline Position wikimedia/wikipedia → this repo →… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-wikipedia.texttext-generation100K<n<1M0 likes49 downloads6mo agoHugging Face07ihalage /sinhala-instruction-finetune-large Dataset Card for sinhala-instruction-finetune-large Sinhala instruction finetune (SIF) dataset contains high quality question-answer pairs in Sinhala language. It is an aggregate of translated English datasets using Google Translate API and several Sinhala datasets in the Hugging Face Datasets hub. SIF dataset has been compiled by transforming the datasets specified below into a common format. sinhala_eli5 sinhala-llm-dataset-llama-prompt-format alpaca-sinhala… See the full description on the dataset page: https://huggingface.co/datasets/ihalage/sinhala-instruction-finetune-large.textquestion-answering100K<n<1M2 likes48 downloads9mo agoHugging Face08Programmer-RD-AI /sinhala-english-singlish-translation Sinhala–English–Singlish Translation Dataset A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations. 📋 Table of Contents Dataset Overview Installation Quick Start Dataset Structure Usage Examples Citation License Credits Dataset Overview Description: 34,500 aligned triplets of Sinhala (native script) English (human translation) Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sinhala-english-singlish-translation.texttranslation10K<n<100K3 likes48 downloads1y agoHugging Face09LexiconShiftInnovations /SinhalaDentalQnAtextquestion-answeringn<1K1 likes44 downloads3y agoHugging Face10Minuri /sinhala-sft-dataset Sinhala Supervised Fine-Tuning Dataset A merged Sinhala instruction-following dataset of 213,703 pairs, used for Supervised Fine-Tuning (SFT) of continually pretrained LLaMA 3.2 1B variants. Constructed as part of a diversity-driven Sinhala language model adaptation study. Dataset Description This dataset merges three existing Sinhala instruction datasets into a unified resource for SFT. It follows the standard Alpaca-style instruction–input–output format and covers a… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-sft-dataset.texttext-generation100K<n<1M0 likes44 downloads6mo agoHugging Face11ChamaraVishwajithRajapaksha /sinhala-text-dataset Sinhala Continuous Pretraining Corpus Curated by HelaAI Dataset Summary This dataset is a Sinhala-language text corpus assembled for continuous pretraining of language models. It combines multiple sources into a single, cleaned, block-structured corpus: News articles — Sinhala news text extracted from the article_sinhala field of Hamza-Ziyard/CNN-Daily-Mail-Sinhala. O/L Sinhala Buddhism — Sinhala-medium educational text covering the GCE Ordinary Level (O/L)… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/sinhala-text-dataset.texttext-generation1M<n<10M0 likes42 downloads28d agoHugging Face12dimuthulk /sinhala-detoxification-dimuthu-subset 📊 Dataset Card for Sinhala Detoxification - Dimuthu's Subset 📝 Dataset Description Repository: dimuthulk/sinhala-detoxification-dimuthu-subset Language(s) (NLP): Sinhala (si) License: Apache 2.0 📋 Dataset Summary This dataset represents the individual data collection and curation contribution of Dimuthu Rathnayaka for the overarching research project, "Sinhala Offensive Text Detoxification Pipeline". It was developed as part of the academic… See the full description on the dataset page: https://huggingface.co/datasets/dimuthulk/sinhala-detoxification-dimuthu-subset.texttext-generation10K<n<100K0 likes40 downloads26d agoHugging Face13dimuthulk /sinhala-text-detoxification 📊 Dataset Card for Sinhala Text Detoxification Dataset 📝 Dataset Description Repository: dimuthulk/sinhala-text-detoxification Language(s) (NLP): Sinhala (si) License: Apache 2.0 📋 Dataset Summary This dataset serves as the primary generative dataset for the second phase of the "Sinhala Offensive Text Detoxification Pipeline" research project conducted at the University of Kelaniya (Electronics and computer science degree program). It contains… See the full description on the dataset page: https://huggingface.co/datasets/dimuthulk/sinhala-text-detoxification.texttext-generation10K<n<100K0 likes39 downloads26d agoHugging Face14Minuri /diverse_sinhala_dataset Diverse Sinhala Dataset A large-scale, cleaned, deduplicated, and domain-classified Sinhala text corpus compiled for continual pretraining of large language models. Constructed as part of a diversity-driven Sinhala language model adaptation study. This repository serves as pipeline storage for the full corpus construction process, from merged cleaned sentences through to the final high-confidence domain-classified corpus. Files File Rows Columns Description… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/diverse_sinhala_dataset.text-generation10M<n<100M0 likes38 downloads6mo agoHugging Face15Minuri /sinhala-corpus-madlad400 Sinhala Raw Sentences - MADLAD-400 Raw Sinhala sentences extracted and sentence-split from the allenai/MADLAD-400 dataset. This is an intermediate dataset used in the construction of Minuri/diverse_sinhala_dataset. Dataset Structure Column Description text Raw Sinhala sentence source Source identifier (madlad) Split Rows train 7,281,026 Pipeline Position allenai/MADLAD-400 → this repo → Minuri/madlad_cleaned_version →… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-madlad400.texttext-generation1M<n<10M0 likes37 downloads6mo agoHugging Face16Minuri /sinhala-corpus-culturax Sinhala Raw Sentences - CulturaX Raw Sinhala sentences extracted and sentence-split from the uonlp/CulturaX dataset. This is an intermediate dataset used in the construction of Minuri/diverse_sinhala_dataset. Dataset Structure Column Description text Raw Sinhala sentence source Source identifier (culturax) Split Rows train 4,707,451 Pipeline Position uonlp/CulturaX → this repo → Minuri/culturax_cleaned_version →… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-culturax.texttext-generation1M<n<10M0 likes37 downloads6mo agoHugging Face17Minuri /sinhala-test-set-50k Sinhala Test Set - 50K Sentences A held-out Sinhala test set of 50,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for perplexity evaluation of three continually pretrained LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study. Dataset Description This test set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased perplexity measurement. It covers multiple… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-test-set-50k.tabulartext-generation10K<n<100K0 likes36 downloads6mo agoHugging Face18SPEAK-PP /openslr-sinhala-synthetic-spell-errors-quarter Sinhala Dyslexic Spelling Correction Dataset Dataset Description This dataset contains Sinhala and code-mixed (Sinhala-English) text pairs for training spelling correction models, specifically designed to address dyslexia-like spelling errors. Features dyslexic_sentence: Input text with dyslexia-like spelling errors (string) correct_sentence: Corrected output text (string) Dataset Statistics Split Samples Train 37,056 Test 9,265… See the full description on the dataset page: https://huggingface.co/datasets/SPEAK-PP/openslr-sinhala-synthetic-spell-errors-quarter.texttext-generation10K<n<100K0 likes32 downloads8mo agoHugging Face19Minuri /sinhala-corpus-b-random-1m Randomly Curated Sinhala Corpus A randomly sampled subset of 1M Sinhala sentences from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model B) as part of a diversity-driven Sinhala language model adaptation study. Corpus variants in this series: Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline) Minuri/sinhala-corpus-b-random-1m - Random subset (random baseline) - this repo Minuri/sinhala-corpus-c-diverse-1m… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-b-random-1m.tabulartext-generation1M<n<10M0 likes32 downloads6mo agoHugging Face20ChamaraVishwajithRajapaksha /cnn-dailymail-sinhala-continuous-pretrain CNN DailyMail Sinhala Continuous Pretraining Dataset Dataset Description This dataset is designed for continuous pretraining of Sinhala Small Language Models (SLMs) and Large Language Models (LLMs). The dataset was created by processing the original Sinhala news articles from: CNN Daily Mail Sinhala Dataset The article_sinhala field from the original dataset was extracted, cleaned, and concatenated into larger continuous text blocks suitable for language model… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/cnn-dailymail-sinhala-continuous-pretrain.texttext-generationn<1K0 likes32 downloads5mo agoHugging Face21LexiconShiftInnovations /SinhalaCorpusLargetexttext-generation10M<n<100M2 likes31 downloads3y agoHugging Face22Chamaka8 /Serendip-sft-sinhala Serendip-SFT-Sinhala Dataset 🇱🇰 📊 Dataset Summary Serendip-SFT-Sinhala is a large-scale Sinhala instruction-tuning dataset with 293,613 high-quality examples for supervised fine-tuning (SFT) of large language models. Created to train SerendipLLM, a Sinhala language model designed to excel at instruction-following, question-answering, summarization, and text classification. 🌟 Highlights 🇱🇰 293,613 Sinhala examples (largest Sinhala SFT dataset) 📚 4 task… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/Serendip-sft-sinhala.texttext-generation100K<n<1M0 likes31 downloads7mo agoHugging Face23KanishkaRandunu /SinhalaWikipediaArticlestexttext-generation10K<n<100K0 likes28 downloads3y agoHugging Face24AyeshaKalpani98 /Questions_Answers_In_Sinhala_Language@misc{AyeshaKalpani_2024, title={Questions_Answers_In_Sinhala_Language}, author={Ayesha Kalpani}, year={2024}, url={}, } Questions_Answers_In_Sinhala_Language Dataset Description A dataset containing questions and answers in the Sinhala language. This dataset is intended for training and evaluating question-answering models in Sinhala. Dataset Details License This dataset is licensed under the MIT License. Task… See the full description on the dataset page: https://huggingface.co/datasets/AyeshaKalpani98/Questions_Answers_In_Sinhala_Language.textquestion-answeringn<1K0 likes27 downloads2y agoHugging Face25Minuri /sinhala-corpus-a-news-1m News-Only Sinhala Corpus A news-domain subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model A) as part of a diversity-driven Sinhala language model adaptation study at the Informatics Institute of Technology (IIT), Colombo, affiliated with Robert Gordon University (RGU). Corpus variants in this series: Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline) - this repo… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-a-news-1m.tabulartext-generation1M<n<10M0 likes27 downloads6mo agoHugging Face26LexiconShiftInnovations /SinhalaWikipediaArticlestexttext-generation10K<n<100K1 likes24 downloads3y agoHugging Face27sh4lu-z /Sinhala-Mega-Corpus-v1 Sinhala Mega Corpus v1 Description English: Sinhala Mega Corpus v1 is a large-scale, high-quality merged dataset specifically designed for training Sinhala Large Language Models (LLMs) and Tokenizers. It combines several major open-source datasets into a single, unified format, providing a diverse range of linguistic patterns from web crawls, encyclopedic knowledge, and conversational data. සිංහල: Sinhala Mega Corpus v1 යනු සිංහල Large Language Models (LLM) සහ Tokenizers… See the full description on the dataset page: https://huggingface.co/datasets/sh4lu-z/Sinhala-Mega-Corpus-v1.texttext-generation100K<n<1M0 likes24 downloads7mo agoHugging Face28janani-rane /Sinhala-News-Wiki-text-corpus Sinhala-News-Wiki-Text-Corpus Containing news articles from various Sinhala news sites along with Sinhala Wikipedia pages. Dataset Overview Language: Sinhala (සිංහල) Content: Sinhala news articles from various sites Data format: Parquet Number of Records: 18,201 rows (as per current size) Dataset Structure Each record consists of the following fields: category: The news category (e.g., "Other-news, Local-news, wiki, International-news"). site: The site's… See the full description on the dataset page: https://huggingface.co/datasets/janani-rane/Sinhala-News-Wiki-text-corpus.tabulartext-classification10K<n<100K0 likes23 downloads2y agoHugging Face29Navanjana /sinhala-articles Sinhala Articles Dataset A large-scale, high-quality Sinhala text corpus curated from diverse sources including news articles, Wikipedia entries, and general web content. This dataset is designed to support a wide range of Sinhala Natural Language Processing (NLP) tasks. 📊 Dataset Overview Name: Navanjana/sinhala-articles Total Samples: 2,148,688 Languages: Sinhala (si) Features: text: A single column containing Sinhala text passages. Size: Approximately 1M < n <… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/sinhala-articles.texttext-generation1M<n<10M1 likes23 downloads1y agoHugging Face30ihalage /sinhala-finetune-qa-eli5 Dataset Card for sinhala-finetune-qa-eli5 Sinhala question answering (QA) dataset contains a subset of the translated eli5 (explain like I'm 5) English dataset. eli5 is a crowdsourced dataset based mainly on the content from the subreddit r/explainlikeimfive. This is a forum where users post complex questions and other users provide simplified explanations. A subset of eli5 dataset (10k samples) has been machine translated to Sinhala language using the Google Cloud Translation API.… See the full description on the dataset page: https://huggingface.co/datasets/ihalage/sinhala-finetune-qa-eli5.textquestion-answering10K<n<100K2 likes18 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.