CoolFace
Datasetpublic

Zyroxx66/somali-master-pretraining-corpus

πŸ‡ΈπŸ‡΄ Somali Master Pretraining Corpus (176.5k Rows) The Somali Master Pretraining Corpus is a curated, balanced dataset designed for Continued Pre-Training (CPT) and foundational pre-training of Large Language Models (LLMs) in the Somali language (Af-Soomaali). It addresses the fundamental challenges of low-resource NLP for Somali by combining quality-filtered web knowledge, structured modern domain knowledge, and synthetic narrative intelligence (TinyStories). πŸŽ―β€¦ See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/somali-master-pretraining-corpus.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes34downloads
3 commits on main
1d1e1b71mo ago

Upload README.md with huggingface_hub

Zyroxx66
a20865c1mo ago

Upload dataset

Zyroxx66
c385f511mo ago

initial commit

Zyroxx66