CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
010xAIT /sinhala-flantext10M<n<100M3 likes3.1k downloads2y agoHugging Face02NimanthaPerera /sinhala_dataset_vtextn<1K0 likes456 downloads13h agoHugging Face03ChamaraVishwajithRajapaksha /sinhala-22gb-cleaned-datasettext1M<n<10M0 likes428 downloads5mo agoHugging Face049wimu9 /sinhala_dataset_sanitizedtext1K<n<10K0 likes354 downloads3y agoHugging Face05ChamathEka /mini-sinhala-flantabular100K<n<1M6 likes321 downloads2y agoHugging Face06akpsahan /SinhalaCXtext100M<n<1B1 likes313 downloads4mo agoHugging Face07sinhala-nlp /SOLD SOLD - A Benchmark for Sinhala Offensive Language Identification In this repository, we introduce the {S}inhala {O}ffensive {L}anguage {D}ataset (SOLD) and present multiple experiments on this dataset. SOLD is a manually annotated dataset containing 10,000 posts from Twitter annotated as offensive and not offensive at both sentence-level and token-level. SOLD is the largest offensive language dataset compiled for Sinhala. We also introduce SemiSOLD, a larger dataset containing more… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/SOLD.texttext-classification10K<n<100K2 likes292 downloads3y agoHugging Face08Isuru0x01 /sinhala_storiestext10M<n<100M0 likes241 downloads1d agoHugging Face09SPEAK-ASR /openslr-sinhala-asraudio100K<n<1M2 likes228 downloads7mo agoHugging Face10edifier99 /sinhala-openslr-111haudio100K<n<1M0 likes209 downloads4mo agoHugging Face119wimu9 /sinhala_sentences_rawtext1K<n<10K1 likes201 downloads3y agoHugging Face12SPEAK-ASR /openslr-sinhala-asr-normaudio10K<n<100K0 likes201 downloads7mo agoHugging Face13irudachirath /large-sinhala-asr-datasetaudio100K<n<1M1 likes193 downloads1y agoHugging Face149wimu9 /sinhala_30m Dataset Card for "sinhala_30m" More Information needed text10M<n<100M1 likes190 downloads3y agoHugging Face15SPEAK-ASR /openslr-sinhala-asr-depricated-versionaudio100K<n<1M0 likes160 downloads9mo agoHugging Face16outlawmold /sinhala-tts-dataset-archive-20260429-082457 Sinhala TTS Dataset Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare. Stats Metric Value Utterances 218 Train 208 Val 10 Hours 0.51 Mean duration 8.5s Sample rate 22050 Hz Pipeline Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 -> Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB) Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.audiotext-to-speechn<1K0 likes156 downloads5mo agoHugging Face17avishadilhara /sinhala-ocr-lk-acts-1010 🇱🇰 Sinhala OCR - Sri Lankan Acts Dataset Dataset Description This dataset contains 1,010 scanned document images of Sri Lankan legal acts (1980s-2010s) in Sinhala language with ground truth text annotations for Optical Character Recognition (OCR) training and evaluation. Key Features ✅ High-quality scanned document images ✅ Professionally corrected ground truth text ✅ Year-wise metadata for temporal analysis ✅ Pre-split into train/eval/test sets… See the full description on the dataset page: https://huggingface.co/datasets/avishadilhara/sinhala-ocr-lk-acts-1010.imageimage-to-text1K<n<10K1 likes149 downloads3mo agoHugging Face18mteb /SinhalaNewsClassification SinhalaNewsClassification An MTEB dataset Massive Text Embedding Benchmark This file contains news texts (sentences) belonging to 5 different news categories (political, business, technology, sports and Entertainment). The original dataset was released by Nisansa de Silva (Sinhala Text Classification: Observations from the Perspective of a Resource Poor Language, 2015). Task category t2c Domains News, Written Reference… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SinhalaNewsClassification.texttext-classification1K<n<10K0 likes135 downloads1y agoHugging Face19Ransaka /SinhalaASR-testaudio10K<n<100K1 likes114 downloads3y agoHugging Face20sh4lu-z /awesome-dataset-sinhala Mixed Sinhala Dataset (1M+ Rows) | මිශ්‍ර සිංහල දත්ත කට්ටලය (Please find the English description below the Sinhala description) 🇬🇧 English This is a comprehensive dataset containing over one million rows of Sinhala text data. It is highly suitable for training Artificial Intelligence (AI) models and conducting Natural Language Processing (NLP) research. Dataset Details Language: Sinhala (si) Total Rows: 1,079,909 Format: Parquet (Optimized for Hugging… See the full description on the dataset page: https://huggingface.co/datasets/sh4lu-z/awesome-dataset-sinhala.texttext-generation1M<n<10M1 likes97 downloads7mo agoHugging Face21Ransaka /SinhalaASR-1000audio1K<n<10K2 likes95 downloads3y agoHugging Face22edifier99 /sinhala_synthetic_ocr_news_largeCreated from https://www.kaggle.com/code/ransakaravihara/sinhala-ocr-image-creation. To cite the dataset @misc{ransaka_r._2026, author = { Ransaka R. }, title = { sinhala_synthetic_ocr_news_large (Revision bc52307) }, year = 2026, url = { https://huggingface.co/datasets/edifier99/sinhala_synthetic_ocr_news_large }, doi = { 10.57967/hf/9748 }, publisher = { Hugging Face } } image1K<n<10K1 likes90 downloads1mo agoHugging Face239wimu9 /sinhala_dataset_59m Dataset Card for "sinhala_dataset_59m" More Information needed text10M<n<100M3 likes89 downloads3y agoHugging Face249wimu9 /sentencified_v1_sinhala_30m Dataset Card for "sentencified_v1_sinhala_30m" More Information needed text10M<n<100M0 likes83 downloads3y agoHugging Face25Chamaka8 /serendip-cpt-sinhala Serendib LLM CPT Sinhala Corpus A large-scale, deduplicated, quality-filtered Sinhala plain-text corpus built for Continual Pre-Training (CPT) of large language models. This dataset was used to adapt Meta-LLaMA-3-8B to the Sinhala language domain as part of the Serendib LLM Honours Degree Research Project at the University of Central Lancashire (UCLan), 2025–2026. This is one of the largest openly published Sinhala NLP corpora available, containing 23,449,223 training documents… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/serendip-cpt-sinhala.texttext-generation10M<n<100M0 likes79 downloads6mo agoHugging Face26sinhala-nlp /named-entity-recognitiontext1K<n<10K2 likes78 downloads2y agoHugging Face27sinhala-nlp /NSINA-Categoriesgated Sinhala News Category Prediction This is a text classification task created with the NSINA dataset. This dataset is also released with the same license as NSINA. Data Data can be loaded into pandas dataframes using the following code. from datasets import Dataset from datasets import load_dataset train = Dataset.to_pandas(load_dataset('sinhala-nlp/NSINA-Categories', split='train')) test = Dataset.to_pandas(load_dataset('sinhala-nlp/NSINA-Categories', split='test'))… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/NSINA-Categories.texttext-classification10K<n<100K1 likes69 downloads3y agoHugging Face28NLPC-UOM /Sinhala-News-Category-classificationThis file contains news texts (sentences) belonging to 5 different news categories (political, business, technology, sports and Entertainment). The original dataset was released by Nisansa de Silva (Sinhala Text Classification: Observations from the Perspective of a Resource Poor Language, 2015). The original dataset is processed and cleaned of single word texts, English only sentences etc. If you use this dataset, please cite {Nisansa de Silva, Sinhala Text Classification: Observations from… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Category-classification.texttext-classification1K<n<10K2 likes62 downloads4y agoHugging Face29deshanksuman /Augmented_SinhalatoRomanizedSinhala_Dataset Sinhala Romanized Dataset This dataset contains Sinhala text with romanized (transliterated) versions, created using publicly available Sinhala data sources. Dataset Description The Augmented Sinhala to Romanized Sinhala Dataset provides paired examples of Sinhala text and their corresponding romanized transliterations. This dataset aims to facilitate research in Sinhala language processing, particularly for applications that require romanized representations of Sinhala… See the full description on the dataset page: https://huggingface.co/datasets/deshanksuman/Augmented_SinhalatoRomanizedSinhala_Dataset.texttranslation1M<n<10M0 likes62 downloads1y agoHugging Face30sinhala-nlp /SiMTEB-NHPtext1K<n<10K0 likes58 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.