LocaleNLP/AfriCorpus-v1
AfriCorpus v1 AfriCorpus-v1 is the first public release of LocaleNLP's audited, deduplicated, and quality-filtered African language corpus. Built to power the AfriLION LLM project, this dataset directly addresses the Tokenizer Fertility problem that causes all current LLMs to underperform on African languages. Key Statistics Language Code Script CC-100 Source Status Wolof wo Latin CC-100 Audited Swahili sw Latin CC-100 Audited Hausa ha Latin + Ajami… See the full description on the dataset page: https://huggingface.co/datasets/LocaleNLP/AfriCorpus-v1.
013
feat(dataset): add full AfriCorpus-v1 dataset card with QA pipeline and design decisions
initial commit
