datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sinhala-corpus-c-diverse-1m
Diversity-Optimized Sinhala Corpus
A diversity-optimized subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model C) as part of a diversity-driven Sinhala language model adaptation study.
Corpus variants in this series:
Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline)
Minuri/sinhala-corpus-b-random-1m - Random subset (random baseline)… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-c-diverse-1m.sinhala-test-set-50k
Sinhala Test Set - 50K Sentences
A held-out Sinhala test set of 50,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for perplexity evaluation of three continually pretrained LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study.
Dataset Description
This test set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased perplexity measurement. It covers multiple… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-test-set-50k.sinhala-corpus-b-random-1m
Randomly Curated Sinhala Corpus
A randomly sampled subset of 1M Sinhala sentences from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model B) as part of a diversity-driven Sinhala language model adaptation study.
Corpus variants in this series:
Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline)
Minuri/sinhala-corpus-b-random-1m - Random subset (random baseline) - this repo
Minuri/sinhala-corpus-c-diverse-1m… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-b-random-1m.sinhala-corpus-a-news-1m
News-Only Sinhala Corpus
A news-domain subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model A) as part of a diversity-driven Sinhala language model adaptation study at the Informatics Institute of Technology (IIT), Colombo, affiliated with Robert Gordon University (RGU).
Corpus variants in this series:
Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline) - this repo… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-a-news-1m.Sinhala-News-Wiki-text-corpus
Sinhala-News-Wiki-Text-Corpus
Containing news articles from various Sinhala news sites along with Sinhala Wikipedia pages.
Dataset Overview
Language: Sinhala (සිංහල)
Content: Sinhala news articles from various sites
Data format: Parquet
Number of Records: 18,201 rows (as per current size)
Dataset Structure
Each record consists of the following fields:
category: The news category (e.g., "Other-news, Local-news, wiki, International-news").
site: The site's… See the full description on the dataset page: https://huggingface.co/datasets/janani-rane/Sinhala-News-Wiki-text-corpus.sinhala-poems-v1
Sinhala Poems (Filtered)
Curated Sinhala poem blocks extracted from web blogs using a verse-shaped heuristic (v3.4) with weak attributes (theme/mood/style) and stats.
Columns
text: full original block (cleaned)
snippet: first stanza or 8 lines
url: source URL
subtype: poem | song_like | promo_like | unknown
keep: boolean accepted by filter
theme, mood, style, length_class: weak labels
ps_*: structure stats (floats)
Filtering summary
Boilerplate/HTML… See the full description on the dataset page: https://huggingface.co/datasets/manthilaffs/sinhala-poems-v1.sinhala-personas-lk-v0.2-gemini-1000-preview
Sinhala-Personas-LK v0.2 Gemini 1000 Preview
Sinhala-Personas-LK is a preview dataset of fully synthetic Sinhala persona records for Sri Lankan NLP research and evaluation.
Version
0.1-preview
Records
1000 synthetic records.
Language
Sinhala (si).
Country context
Sri Lanka (LK).
Important limitations
This preview version is generated from starter priors and LLM-generated text. It is not yet fully grounded… See the full description on the dataset page: https://huggingface.co/datasets/sayururehan/sinhala-personas-lk-v0.2-gemini-1000-preview.sinhala-validation-set-10k
Sinhala Validation Set - 10K Sentences
A held-out Sinhala validation set of 10,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for monitoring validation loss during continual pretraining of three LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study.
Dataset Description
This validation set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased validation loss… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-validation-set-10k.
