datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SemiSOLD
SOLD - A Benchmark for Sinhala Offensive Language Identification
In this repository, we introduce the {S}inhala {O}ffensive {L}anguage {D}ataset (SOLD) and present multiple experiments on this dataset. SOLD is a manually annotated dataset containing 10,000 posts from Twitter annotated as offensive and not offensive at both sentence-level and token-level. SOLD is the largest offensive language dataset compiled for Sinhala. We also introduce SemiSOLD, a larger dataset containing more… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/SemiSOLD.Sinhala-News-Source-classificationThis dataset contains Sinhala news headlines extracted from 9 news sources (websites) (Sri Lanka Army, Dinamina, GossipLanka, Hiru, ITN, Lankapuwath, NewsLK,
Newsfirst, World Socialist Web Site-Sinhala). This is a processed version of the corpus created by Sachintha, D., Piyarathna, L., Rajitha, C., and Ranathunga, S. (2021). Exploiting parallel corpora to improve multilingual embedding based document and sentence alignment. Single word sentences, invalid characters have been removed from the… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Source-classification.sinhala-summarization-dataset
Sinhala Text Summarization Dataset
Dataset Description
This dataset is a Sinhala text summarization dataset created for research in low-resource language summarization. The dataset contains 2,493 Sinhala article-summary pairs collected from diverse publicly accessible Sinhala online sources.
This repository contains a Sinhala article-summary dataset introduced in the following IEEE conference publication:
Sinhala Automatic Text Summarization: Dataset Creation and… See the full description on the dataset page: https://huggingface.co/datasets/hans1k/sinhala-summarization-dataset.akura-sinhala-dyslexic-writing-patterns
Akura Sinhala Dyslexic Writing Patterns Dataset
Overview
This dataset provides a sentence-level, feature-augmented corpus for diagnosing dyslexic writing patterns in Sinhala.Each instance consists of a dyslexic sentence, its corresponding clean reference sentence, a set of explicit character-level error features, and a dominant dyslexic writing pattern label.
Unlike correction-focused datasets, this corpus is designed for diagnostic classification, enabling models to… See the full description on the dataset page: https://huggingface.co/datasets/akura-official/akura-sinhala-dyslexic-writing-patterns.FacebookDecadeCorporasinhala_YouTube_Comment_Processed
