CoolFace
Datasetpublic

mamakobe/luhya-multilingual-dataset

Luhya Multilingual Dataset Curator: Dr. Moody AmakobeProject: Project Tafsiri — Bridging Indigenous Languages and AIVersion: 2.0License: Creative Commons Attribution 4.0 (CC BY 4.0) Dataset Overview This dataset is a comprehensive, structured multilingual corpus for the Luhya language (also written Luyia), a Bantu language cluster spoken primarily in western Kenya by the Abaluhya people — Kenya's second largest ethnic group with… See the full description on the dataset page: https://huggingface.co/datasets/mamakobe/luhya-multilingual-dataset.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
4likes47downloads
Dataset Card

Luhya Multilingual Dataset

Curator: Dr. Moody Amakobe Project: Project Tafsiri — Bridging Indigenous Languages and AI Version: 2.0 License: Creative Commons Attribution 4.0 (CC BY 4.0)


Dataset Overview

This dataset is a comprehensive, structured multilingual corpus for the Luhya language (also written Luyia), a Bantu language cluster spoken primarily in western Kenya by the Abaluhya people — Kenya's second largest ethnic group with approximately 7 million speakers.

The dataset was created as part of Project Tafsiri, an initiative to bridge indigenous Kenyan languages with modern AI systems by building high-quality, culturally grounded training data for low-resource language models.

The corpus covers 4 domains, 9 dialects, and provides trilingual alignment across Luhya ↔ English ↔ Swahili — making it suitable for machine translation, language modelling, dialect classification, and cross-lingual transfer learning tasks.


Dataset Statistics

MetricCount
Total rows22,269
Dialects covered9
Domains covered4
Rows with English22,268 (99.9%)
Rows with Swahili6,534 (29.3%)
Fully trilingual rows6,533

Domains

DomainRowsDescription
Dictionary11,493Word-level entries with English definitions, part of speech, and notes
Conversational6,447Naturally occurring sentences from community speech
Proverbs2,331Traditional Luhya proverbs with cultural annotations
Bible1,998Parallel scripture verses (Luhya ↔ English)

Dialects Covered

The Luhya language is a cluster of 16–18 mutually intelligible dialects. This dataset currently covers the following:

DialectDescription
BukusuWidely spoken in Bungoma County; famous for circumcision ceremonies and oral literature
Wanga / LuwangaKakamega region dialect with rich oral traditions
MaragoliSpoken in Vihiga County; known for formal writing and traditional music
MarachiSpoken by the Marachi people in Butula, Busia County
KisaClose to Idakho; spoken in Kakamega County
MaramaCentral Luhya cluster, Kakamega County
TsotsoClosely related to other central dialects, Kakamega County
IsukhaKakamega region; closely related to Idakho
NyalaSpoken by the Nyala people in Busia district

Dataset Fields

FieldTypeDescription
idstringUnique UUID for each row
luhya_textstringSource text in Luhya (always present)
english_textstringEnglish translation or definition (99.9% coverage)
swahili_textstringSwahili translation (29% coverage, primarily conversational domain)
dialect_idstringUUID linking to the dialect reference table
dialect_namestringHuman-readable dialect name (e.g. "Bukusu", "Wanga")
domainstringContent domain: Bible / Dictionary / Proverbs / Conversational
subdomainstringFiner-grained context (e.g. "Genesis 1:1", "Bukusu Dictionary", "general_wisdom")
posstringPart of speech — Dictionary rows only (n, v, adj, adv, etc.)
quality_scorefloatQuality score where available (Bible domain)
is_validatedboolWhether the entry has been human-validated
notesstringCultural context, noun class, cross-references, etymology
source_tablestringOriginating database table for provenance tracing
source_idstringRow ID in the originating table

Data Sources

This dataset was compiled from the following primary sources:

1. Lubukusu-English Dictionary (Academic)

Marlo, Michael; Sifuna, Adrian; Wasike, Aggrey (2008) A comprehensive ~6,200-word Bukusu-English dictionary compiled at Indiana University and Michigan State University. Peer-reviewed and tone-marked. Available on Academia.edu.

2. Wanga-English Dictionary (Academic)

Compiled with Alfred Anangwe (2008) A ~4,000-word preliminary Wanga-English dictionary compiled as part of the Luyia Dictionary Project at the University of Michigan. Available on Academia.edu.

3. IBIBILIA INDAKATIFU — The New Testament in Luhya

Matayo–Obufwimbuli | Oluluyia Bible Society of Kenya, 2019 Provides formal religious language structure and vocabulary. A foundational text for standardised Luhya orthography.

4. KenTrans: A Parallel Corpora for Swahili and Local Kenyan Languages

Wanzare, Lilian D.A; Indede, Florence; McOnyango, Owen; Ombui, Edward; Wanjawa, Barack; Muchemi, Lawrence (2022) https://doi.org/10.7910/DVN/NOAT0W — Harvard Dataverse, V2 Offers professionally validated translation pairs and linguistic patterns across Swahili and local Kenyan languages including Luhya.

5. Collection of 100 Nyala Proverbs

Kevin Namatsi Okubo Africa Proverb Working Group, Nairobi, Kenya, April 2016 Preserves traditional wisdom and cultural knowledge systems specific to the Nyala dialect.

6. Luyia Proverbs from Kisa, Marama, Tsotso and Wanga

Tim Wambunya Provides cross-dialectal cultural content and variation patterns across four Luhya dialects.

7. Community Dictionary Contributions

Structured word entries collected from native speakers across Wanga and Luwanga dialects, including noun class information, etymology, pronunciation notes, and cultural context.


About Project Tafsiri

Project Tafsiri is an initiative dedicated to building AI infrastructure for indigenous Kenyan languages. The project focuses on creating high-quality, culturally authentic training data to ensure that AI systems can serve speakers of low-resource languages — particularly those in communities that have historically been excluded from the benefits of language technology.

The LuhyaAI system built on this dataset is grounded in both contemporary usage and traditional knowledge. By combining conversational sentences, formal scripture, academic dictionaries, and community proverbs, the dataset captures the full linguistic and cultural range of the Luhya language — creating a foundation for AI systems that are not only accurate but culturally resonant.


Citation

If you use this dataset in your research, please cite:

@dataset{amakobe2025luhya,
  author    = {Amakobe, Moody},
  title     = {Luhya Multilingual Dataset},
  year      = {2025},
  publisher = {HuggingFace},
  url       = {https://huggingface.co/datasets/mamakobe/luhya-multilingual-dataset},
  note      = {Part of Project Tafsiri — https://www.riskinfo.ai/post/project-tafsiri-bridging-indigenous-languages-and-ai}
}

Contact

For questions, contributions, or collaboration opportunities related to Project Tafsiri or the LuhyaAI initiative, please reach out via the Project Tafsiri page.