CoolFace
Datasetpublic

LocaleNLP/AfriCorpus-v1

AfriCorpus v1 AfriCorpus-v1 is the first public release of LocaleNLP's audited, deduplicated, and quality-filtered African language corpus. Built to power the AfriLION LLM project, this dataset directly addresses the Tokenizer Fertility problem that causes all current LLMs to underperform on African languages. Key Statistics Language Code Script CC-100 Source Status Wolof wo Latin CC-100 Audited Swahili sw Latin CC-100 Audited Hausa ha Latin + Ajami… See the full description on the dataset page: https://huggingface.co/datasets/LocaleNLP/AfriCorpus-v1.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes13downloads
2 commits on main
1022b095mo ago

feat(dataset): add full AfriCorpus-v1 dataset card with QA pipeline and design decisions

aljagne
718006e5mo ago

initial commit

aljagne