LocaleNLP/AfriCorpus-v1
AfriCorpus v1 AfriCorpus-v1 is the first public release of LocaleNLP's audited, deduplicated, and quality-filtered African language corpus. Built to power the AfriLION LLM project, this dataset directly addresses the Tokenizer Fertility problem that causes all current LLMs to underperform on African languages. Key Statistics Language Code Script CC-100 Source Status Wolof wo Latin CC-100 Audited Swahili sw Latin CC-100 Audited Hausa ha Latin + Ajami… See the full description on the dataset page: https://huggingface.co/datasets/LocaleNLP/AfriCorpus-v1.
This repository belongs to LocaleNLP on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
