Minuri/sinhala-corpus-a-news-1m
News-Only Sinhala Corpus A news-domain subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model A) as part of a diversity-driven Sinhala language model adaptation study at the Informatics Institute of Technology (IIT), Colombo, affiliated with Robert Gordon University (RGU). Corpus variants in this series: Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline) - this… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-a-news-1m.
News-Only Sinhala Corpus
A news-domain subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model A) as part of a diversity-driven Sinhala language model adaptation study at the Informatics Institute of Technology (IIT), Colombo, affiliated with Robert Gordon University (RGU).
Corpus variants in this series: -Minuri/sinhala-corpus-a-news-1m- News-only subset (domain-homogeneous baseline) - this repo -Minuri/sinhala-corpus-b-random-1m- Random subset (random baseline) -Minuri/sinhala-corpus-c-diverse-1m- Diversity-optimized subset ✅ Best perplexity
Dataset Description
Corpus A serves as the domain-homogeneous baseline, comprising sentences drawn exclusively from the news domain of the parent corpus. This enables controlled comparison against the random (B) and diversity-optimized (C) corpora in downstream perplexity and evaluation experiments. The model trained on this corpus (Model A) achieved a perplexity of 14.68 on the Sinhala test set.
Source Datasets (via parent corpus)
Dataset Structure
Splits
Format
Available in both JSONL and CSV formats.
Intended Uses
- Continual pretraining of LLMs on Sinhala (domain-homogeneous baseline)
- Ablation studies on corpus diversity
- Sinhala NLP benchmarking
Associated Model
This corpus was used to train: Minuri/sinhala-llama-1b-corpus-news
Sources & Licenses
This dataset contains sentences derived from the following source datasets. Users must comply with the license terms of each:
This dataset is released under CC BY-SA 4.0 in compliance with the ShareAlike terms of Wikipedia and NSINA.
