ChamaraVishwajithRajapaksha/sinhala-text-dataset
Sinhala Continuous Pretraining Corpus Curated by HelaAI Dataset Summary This dataset is a Sinhala-language text corpus assembled for continuous pretraining of language models. It combines multiple sources into a single, cleaned, block-structured corpus: News articles — Sinhala news text extracted from the article_sinhala field of Hamza-Ziyard/CNN-Daily-Mail-Sinhala. O/L Sinhala Buddhism — Sinhala-medium educational text covering the GCE Ordinary Level (O/L)… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/sinhala-text-dataset.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face