CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ChamathEka /mini-sinhala-flantabular100K<n<1M6 likes321 downloads2y agoHugging Face02Minuri /sinhala-corpus-c-diverse-1m Diversity-Optimized Sinhala Corpus A diversity-optimized subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model C) as part of a diversity-driven Sinhala language model adaptation study. Corpus variants in this series: Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline) Minuri/sinhala-corpus-b-random-1m - Random subset (random baseline)… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-c-diverse-1m.tabulartext-generation1M<n<10M0 likes50 downloads6mo agoHugging Face03sinhala-nlp /SemiSOLD SOLD - A Benchmark for Sinhala Offensive Language Identification In this repository, we introduce the {S}inhala {O}ffensive {L}anguage {D}ataset (SOLD) and present multiple experiments on this dataset. SOLD is a manually annotated dataset containing 10,000 posts from Twitter annotated as offensive and not offensive at both sentence-level and token-level. SOLD is the largest offensive language dataset compiled for Sinhala. We also introduce SemiSOLD, a larger dataset containing more… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/SemiSOLD.tabular100K<n<1M0 likes45 downloads3y agoHugging Face04Siluni /sinhala-vqa-dataset Sinhala VQA Dataset A Sinhala-language Visual Question Answering dataset of 37,318 QA pairs, constructed by translating Visual Genome QA annotations into Sinhala using gemini-3-flash-preview. This dataset was developed as part of research on benchmarking and adapting compact multimodal models for Sinhala VQA under low-resource conditions. Dataset Summary Split Samples Train 33,409 Validation 2,909 Test 1,000 Total 37,318 Schema Each row… See the full description on the dataset page: https://huggingface.co/datasets/Siluni/sinhala-vqa-dataset.tabularvisual-question-answering10K<n<100K0 likes38 downloads6mo agoHugging Face05NLPC-UOM /Sinhala-News-Source-classificationThis dataset contains Sinhala news headlines extracted from 9 news sources (websites) (Sri Lanka Army, Dinamina, GossipLanka, Hiru, ITN, Lankapuwath, NewsLK, Newsfirst, World Socialist Web Site-Sinhala). This is a processed version of the corpus created by Sachintha, D., Piyarathna, L., Rajitha, C., and Ranathunga, S. (2021). Exploiting parallel corpora to improve multilingual embedding based document and sentence alignment. Single word sentences, invalid characters have been removed from the… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Source-classification.tabulartext-classification10K<n<100K0 likes37 downloads4y agoHugging Face06Minuri /sinhala-test-set-50k Sinhala Test Set - 50K Sentences A held-out Sinhala test set of 50,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for perplexity evaluation of three continually pretrained LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study. Dataset Description This test set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased perplexity measurement. It covers multiple… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-test-set-50k.tabulartext-generation10K<n<100K0 likes31 downloads6mo agoHugging Face07hans1k /sinhala-summarization-dataset Sinhala Text Summarization Dataset Dataset Description This dataset is a Sinhala text summarization dataset created for research in low-resource language summarization. The dataset contains 2,493 Sinhala article-summary pairs collected from diverse publicly accessible Sinhala online sources. This repository contains a Sinhala article-summary dataset introduced in the following IEEE conference publication: Sinhala Automatic Text Summarization: Dataset Creation and… See the full description on the dataset page: https://huggingface.co/datasets/hans1k/sinhala-summarization-dataset.tabularsummarization1K<n<10K0 likes28 downloads4mo agoHugging Face08Minuri /sinhala-corpus-b-random-1m Randomly Curated Sinhala Corpus A randomly sampled subset of 1M Sinhala sentences from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model B) as part of a diversity-driven Sinhala language model adaptation study. Corpus variants in this series: Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline) Minuri/sinhala-corpus-b-random-1m - Random subset (random baseline) - this repo Minuri/sinhala-corpus-c-diverse-1m… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-b-random-1m.tabulartext-generation1M<n<10M0 likes26 downloads6mo agoHugging Face09janani-rane /Sinhala-News-Wiki-text-corpus Sinhala-News-Wiki-Text-Corpus Containing news articles from various Sinhala news sites along with Sinhala Wikipedia pages. Dataset Overview Language: Sinhala (සිංහල) Content: Sinhala news articles from various sites Data format: Parquet Number of Records: 18,201 rows (as per current size) Dataset Structure Each record consists of the following fields: category: The news category (e.g., "Other-news, Local-news, wiki, International-news"). site: The site's… See the full description on the dataset page: https://huggingface.co/datasets/janani-rane/Sinhala-News-Wiki-text-corpus.tabulartext-classification10K<n<100K0 likes22 downloads2y agoHugging Face10Minuri /sinhala-corpus-a-news-1m News-Only Sinhala Corpus A news-domain subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model A) as part of a diversity-driven Sinhala language model adaptation study at the Informatics Institute of Technology (IIT), Colombo, affiliated with Robert Gordon University (RGU). Corpus variants in this series: Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline) - this repo… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-a-news-1m.tabulartext-generation1M<n<10M0 likes21 downloads6mo agoHugging Face11naist-nlp /SinhalaMMLUgated SinhalaMMLU We introduce SinhalaMMLU, the first multiple-choice question answering benchmark designed specifically for Sinhala, a low-resource language.The dataset contains over 7,000 questions spanning secondary to collegiate education levels, aligned with the Sri Lankan national curriculum.It covers six domains and 30 subjects, encompassing both general academic topics and culturally grounded knowledge. We evaluated 26 large language models (LLMs) on SinhalaMMLU and observed that… See the full description on the dataset page: https://huggingface.co/datasets/naist-nlp/SinhalaMMLU.tabular1K<n<10K0 likes19 downloads7mo agoHugging Face12sayururehan /sinhala-personas-lk-v0.2-gemini-1000-preview Sinhala-Personas-LK v0.2 Gemini 1000 Preview Sinhala-Personas-LK is a preview dataset of fully synthetic Sinhala persona records for Sri Lankan NLP research and evaluation. Version 0.1-preview Records 1000 synthetic records. Language Sinhala (si). Country context Sri Lanka (LK). Important limitations This preview version is generated from starter priors and LLM-generated text. It is not yet fully grounded… See the full description on the dataset page: https://huggingface.co/datasets/sayururehan/sinhala-personas-lk-v0.2-gemini-1000-preview.tabulartext-generation1K<n<10K0 likes17 downloads4mo agoHugging Face13Lingalingeswaran /Testing_audio_text_pairs_Sinhala_v1tabularn<1K0 likes15 downloads1y agoHugging Face14theekshana /sinhala-news-sentiment-classification Dataset Card for "sinhala-news-sentiment-classification" More Information needed tabular10K<n<100K0 likes14 downloads2y agoHugging Face15manthilaffs /sinhala-poems-v1 Sinhala Poems (Filtered) Curated Sinhala poem blocks extracted from web blogs using a verse-shaped heuristic (v3.4) with weak attributes (theme/mood/style) and stats. Columns text: full original block (cleaned) snippet: first stanza or 8 lines url: source URL subtype: poem | song_like | promo_like | unknown keep: boolean accepted by filter theme, mood, style, length_class: weak labels ps_*: structure stats (floats) Filtering summary Boilerplate/HTML… See the full description on the dataset page: https://huggingface.co/datasets/manthilaffs/sinhala-poems-v1.tabulartext-generationn<1K0 likes13 downloads1y agoHugging Face16haturusinghe /facebook-sinhala-unicode-hate-speechtabular1K<n<10K0 likes11 downloads2y agoHugging Face17ovinduG /gemma4-sinhala-cpt-evaltabularn<1K0 likes11 downloads3mo agoHugging Face18akura-official /akura-sinhala-dyslexic-writing-patterns Akura Sinhala Dyslexic Writing Patterns Dataset Overview This dataset provides a sentence-level, feature-augmented corpus for diagnosing dyslexic writing patterns in Sinhala.Each instance consists of a dyslexic sentence, its corresponding clean reference sentence, a set of explicit character-level error features, and a dominant dyslexic writing pattern label. Unlike correction-focused datasets, this corpus is designed for diagnostic classification, enabling models to… See the full description on the dataset page: https://huggingface.co/datasets/akura-official/akura-sinhala-dyslexic-writing-patterns.tabular10K<n<100K0 likes10 downloads8mo agoHugging Face19Minuri /sinhala-validation-set-10k Sinhala Validation Set - 10K Sentences A held-out Sinhala validation set of 10,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for monitoring validation loss during continual pretraining of three LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study. Dataset Description This validation set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased validation loss… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-validation-set-10k.tabulartext-generation10K<n<100K0 likes10 downloads6mo agoHugging Face20naist-nlp /SinhalaMMLU-Completegated SinhalaMMLU We introduce SinhalaMMLU, the first multiple-choice question answering benchmark designed specifically for Sinhala, a low-resource language.The dataset contains over 7,000 questions spanning secondary to collegiate education levels, aligned with the Sri Lankan national curriculum.It covers six domains and 30 subjects, encompassing both general academic topics and culturally grounded knowledge. We evaluated 26 large language models (LLMs) on SinhalaMMLU and observed… See the full description on the dataset page: https://huggingface.co/datasets/naist-nlp/SinhalaMMLU-Complete.tabular1K<n<10K0 likes10 downloads3mo agoHugging Face21sinhala-nlp /FacebookDecadeCorporatabular100K<n<1M0 likes8 downloads2y agoHugging Face22ov1n /sinhala-alevel-physics-questionsgated Dataset Details This dataset contains 20 physics questions and answers focused on Sinhala language. tabularquestion-answeringn<1K0 likes5 downloads2y agoHugging Face23Yomaldm /sinhala_YouTube_Emotion_Splitstabular100K<n<1M0 likes5 downloads1y agoHugging Face24Achintha0626 /sinhala-news-sentiment-classification Dataset Card for "sinhala-news-sentiment-classification" More Information needed tabular10K<n<100K0 likes4 downloads6mo agoHugging Face25Yomaldm /sinhala_YouTube_Comment_Processedtabular10K<n<100K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.