CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Programmer-RD-AI /sinhala-english-singlish-translation Sinhala–English–Singlish Translation Dataset A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations. 📋 Table of Contents Dataset Overview Installation Quick Start Dataset Structure Usage Examples Citation License Credits Dataset Overview Description: 34,500 aligned triplets of Sinhala (native script) English (human translation) Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sinhala-english-singlish-translation.texttranslation10K<n<100K3 likes47 downloads1y agoHugging Face02Navanjana /sinhala-articles Sinhala Articles Dataset A large-scale, high-quality Sinhala text corpus curated from diverse sources including news articles, Wikipedia entries, and general web content. This dataset is designed to support a wide range of Sinhala Natural Language Processing (NLP) tasks. 📊 Dataset Overview Name: Navanjana/sinhala-articles Total Samples: 2,148,688 Languages: Sinhala (si) Features: text: A single column containing Sinhala text passages. Size: Approximately 1M < n <… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/sinhala-articles.texttext-generation1M<n<10M1 likes18 downloads1y agoHugging Face03lm-spell /sinhala-spell-correction-datasetgated Sinhala Spell Correction Dataset A Sinhala spell correction dataset used for training and evaluating neural spell correction models as part of the LMSpell project. Dataset Description This dataset combines data from previously published Sinhala spell correction resources and applies additional cleaning to improve its suitability for training neural spell correction models. The dataset originates from the benchmark introduced by Sonnadara et al. (2021) and was… See the full description on the dataset page: https://huggingface.co/datasets/lm-spell/sinhala-spell-correction-dataset.texttext-generation100K<n<1M0 likes12 downloads1d agoHugging Face04Navanjana /Sinhala-Wikitexttext-generation10K<n<100K0 likes9 downloads2y agoHugging Face05NLPC-UOM /anonymized-sinhala-letter-corpus Anonymized Sinhala Official Letter Corpus A small, hand-curated corpus of 151 formal Sinhala letters, fully anonymized with bracketed placeholders. It is intended for training and evaluating models that generate, complete, or classify Sinhala official correspondence — a task with very little public training data. Dataset at a glance Examples 151 Language Sinhala (si) Register Formal throughout Letter length 42–240 words (median 108, mean 114)… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/anonymized-sinhala-letter-corpus.texttext-generationn<1K0 likes7 downloads2mo agoHugging Face06sinhala-nlp /NSINA-Headlinesgated Sinhala Headline Generation This is a text generation task created with the NSINA dataset. This dataset is also released with the same license as NSINA. The objective of the task is to generate news headlines based on the provided news content. Data We used the same instances from NSINA 1.0 as all the news articles had headlines. We divided this dataset into a training and test set following a 0.8 split. Data can be loaded into pandas dataframes using the following code.… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/NSINA-Headlines.texttext-generation100K<n<1M0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.