CoolFace
Datasetpublic

Zyroxx66/somali-master-pretraining-corpus

πŸ‡ΈπŸ‡΄ Somali Master Pretraining Corpus (176.5k Rows) The Somali Master Pretraining Corpus is a curated, balanced dataset designed for Continued Pre-Training (CPT) and foundational pre-training of Large Language Models (LLMs) in the Somali language (Af-Soomaali). It addresses the fundamental challenges of low-resource NLP for Somali by combining quality-filtered web knowledge, structured modern domain knowledge, and synthetic narrative intelligence (TinyStories). πŸŽ―β€¦ See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/somali-master-pretraining-corpus.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes42downloads
settings

This repository belongs to Zyroxx66 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namesomali-master-pretraining-corpus
visibilitypublic
licenceapache-2.0
gatedno
ownerZyroxx66
Account settings
Zyroxx66/somali-master-pretraining-corpus Β· CoolFace