datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AfriCorpus-v1
AfriCorpus v1
AfriCorpus-v1 is the first public release of LocaleNLP's audited, deduplicated, and quality-filtered African language corpus. Built to power the AfriLION LLM project, this dataset directly addresses the Tokenizer Fertility problem that causes all current LLMs to underperform on African languages.
Key Statistics
Language
Code
Script
CC-100 Source
Status
Wolof
wo
Latin
CC-100
Audited
Swahili
sw
Latin
CC-100
Audited
Hausa
ha
Latin + Ajami
CC-100… See the full description on the dataset page: https://huggingface.co/datasets/LocaleNLP/AfriCorpus-v1.Multilingual_Locale_Fidelity
🇰🇿 Kazakh Multi-Tool Agentic AI and Locale Fidelity Dataset
Dataset Summary
Kazakh Multi-Tool Agentic AI and Locale Fidelity Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in agentic AI scenarios that require multi-step tool use, function calling, and locale-sensitive reasoning.
The dataset focuses on practical assistant workflows where a model must understand a user request in Kazakh, select the correct… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Multilingual_Locale_Fidelity.
