language-exposure
swahili-language-exposure
swahili-language-exposure
Dataset Summary
swahili-language-exposure is a large-scale Swahili (Kiswahili) corpus designed for language exposure and continued pretraining of language models.
Unlike instruction-tuning datasets, this dataset focuses on exposing models to natural Swahili usage across conversations, explanations, narratives, technical discussions, and mixed-domain text. The goal is to improve fluency, vocabulary coverage, syntax, and cultural grounding in… See the full description on the dataset page: https://huggingface.co/datasets/nileagi/swahili-language-exposure.swahili-language-exposure-v2
Swahili Language Exposure
Large-scale Swahili corpus for continued pretraining and language exposure.
Maintained by NileAGI.
