datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kenyan-languages-corpus-engine
🇰🇪 LughaGen Modular Engine
LughaGen is a high-performance, modular data preprocessing and normalization engine designed to build high-quality parallel corpora for low-resource Kenyan and regional languages.
The pipeline dynamically loads, normalizes, cleans, and partitions raw parallel datasets (CSV, Excel, Parquet, JSONL) into stratified train, validation, and test splits ready for Machine Learning, Neural Machine Translation (NMT), and LLM tokenizer training.… See the full description on the dataset page: https://huggingface.co/datasets/samptah/kenyan-languages-corpus-engine.kenyan-languages-corpus-enginejambogpt-kenyan-languages
JamboGPT Kenyan Languages Voice Dataset
Overview
This is a comprehensive voice dataset for Kenyan languages, created to advance AI accessibility in African languages. The dataset contains high-quality speech recordings with transcriptions for 5 major Kenyan languages.
Languages Included
Language
Code
Speakers
Region
Samples
Swahili
swh
100M+
East Africa
5,000+
Kikuyu
ki
7M
Central Kenya
5,000+
Luo
luo
4M
Western Kenya
5,000+
Luhya
luy
5M… See the full description on the dataset page: https://huggingface.co/datasets/stano03/jambogpt-kenyan-languages.
