CoolFace
Datasetpublic

Minuri/diverse_sinhala_dataset

Diverse Sinhala Dataset A large-scale, cleaned, deduplicated, and domain-classified Sinhala text corpus compiled for continual pretraining of large language models. Constructed as part of a diversity-driven Sinhala language model adaptation study. This repository serves as pipeline storage for the full corpus construction process, from merged cleaned sentences through to the final high-confidence domain-classified corpus. Files File Rows Columns Description… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/diverse_sinhala_dataset.

sourceHugging Facecc-by-sa-4.0updated 6mo agoView on Hugging Face
0likes38downloads
settings

This repository belongs to Minuri on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namediverse_sinhala_dataset
visibilitypublic
licencecc-by-sa-4.0
gatedno
ownerMinuri
Account settings
Minuri/diverse_sinhala_dataset · CoolFace