Minuri/diverse_sinhala_dataset
Diverse Sinhala Dataset A large-scale, cleaned, deduplicated, and domain-classified Sinhala text corpus compiled for continual pretraining of large language models. Constructed as part of a diversity-driven Sinhala language model adaptation study. This repository serves as pipeline storage for the full corpus construction process, from merged cleaned sentences through to the final high-confidence domain-classified corpus. Files File Rows Columns Description… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/diverse_sinhala_dataset.
This repository belongs to Minuri on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
