mamakobe/luhya-multilingual-dataset
Luhya Multilingual Dataset Curator: Dr. Moody AmakobeProject: Project Tafsiri — Bridging Indigenous Languages and AIVersion: 2.0License: Creative Commons Attribution 4.0 (CC BY 4.0) Dataset Overview This dataset is a comprehensive, structured multilingual corpus for the Luhya language (also written Luyia), a Bantu language cluster spoken primarily in western Kenya by the Abaluhya people — Kenya's second largest ethnic group with… See the full description on the dataset page: https://huggingface.co/datasets/mamakobe/luhya-multilingual-dataset.
Luhya Multilingual Dataset
Curator: Dr. Moody Amakobe Project: Project Tafsiri — Bridging Indigenous Languages and AI Version: 2.0 License: Creative Commons Attribution 4.0 (CC BY 4.0)
Dataset Overview
This dataset is a comprehensive, structured multilingual corpus for the Luhya language (also written Luyia), a Bantu language cluster spoken primarily in western Kenya by the Abaluhya people — Kenya's second largest ethnic group with approximately 7 million speakers.
The dataset was created as part of Project Tafsiri, an initiative to bridge indigenous Kenyan languages with modern AI systems by building high-quality, culturally grounded training data for low-resource language models.
The corpus covers 4 domains, 9 dialects, and provides trilingual alignment across Luhya ↔ English ↔ Swahili — making it suitable for machine translation, language modelling, dialect classification, and cross-lingual transfer learning tasks.
Dataset Statistics
Domains
Dialects Covered
The Luhya language is a cluster of 16–18 mutually intelligible dialects. This dataset currently covers the following:
Dataset Fields
Data Sources
This dataset was compiled from the following primary sources:
1. Lubukusu-English Dictionary (Academic)
Marlo, Michael; Sifuna, Adrian; Wasike, Aggrey (2008) A comprehensive ~6,200-word Bukusu-English dictionary compiled at Indiana University and Michigan State University. Peer-reviewed and tone-marked. Available on Academia.edu.
2. Wanga-English Dictionary (Academic)
Compiled with Alfred Anangwe (2008) A ~4,000-word preliminary Wanga-English dictionary compiled as part of the Luyia Dictionary Project at the University of Michigan. Available on Academia.edu.
3. IBIBILIA INDAKATIFU — The New Testament in Luhya
Matayo–Obufwimbuli | Oluluyia Bible Society of Kenya, 2019 Provides formal religious language structure and vocabulary. A foundational text for standardised Luhya orthography.
4. KenTrans: A Parallel Corpora for Swahili and Local Kenyan Languages
Wanzare, Lilian D.A; Indede, Florence; McOnyango, Owen; Ombui, Edward; Wanjawa, Barack; Muchemi, Lawrence (2022) https://doi.org/10.7910/DVN/NOAT0W — Harvard Dataverse, V2 Offers professionally validated translation pairs and linguistic patterns across Swahili and local Kenyan languages including Luhya.
5. Collection of 100 Nyala Proverbs
Kevin Namatsi Okubo Africa Proverb Working Group, Nairobi, Kenya, April 2016 Preserves traditional wisdom and cultural knowledge systems specific to the Nyala dialect.
6. Luyia Proverbs from Kisa, Marama, Tsotso and Wanga
Tim Wambunya Provides cross-dialectal cultural content and variation patterns across four Luhya dialects.
7. Community Dictionary Contributions
Structured word entries collected from native speakers across Wanga and Luwanga dialects, including noun class information, etymology, pronunciation notes, and cultural context.
About Project Tafsiri
Project Tafsiri is an initiative dedicated to building AI infrastructure for indigenous Kenyan languages. The project focuses on creating high-quality, culturally authentic training data to ensure that AI systems can serve speakers of low-resource languages — particularly those in communities that have historically been excluded from the benefits of language technology.
The LuhyaAI system built on this dataset is grounded in both contemporary usage and traditional knowledge. By combining conversational sentences, formal scripture, academic dictionaries, and community proverbs, the dataset captures the full linguistic and cultural range of the Luhya language — creating a foundation for AI systems that are not only accurate but culturally resonant.
Citation
If you use this dataset in your research, please cite:
@dataset{amakobe2025luhya,
author = {Amakobe, Moody},
title = {Luhya Multilingual Dataset},
year = {2025},
publisher = {HuggingFace},
url = {https://huggingface.co/datasets/mamakobe/luhya-multilingual-dataset},
note = {Part of Project Tafsiri — https://www.riskinfo.ai/post/project-tafsiri-bridging-indigenous-languages-and-ai}
}Contact
For questions, contributions, or collaboration opportunities related to Project Tafsiri or the LuhyaAI initiative, please reach out via the Project Tafsiri page.
