datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dholuodholuo-speech-data
Dholuo Speech Data (Pooled)
A ~191.5-hour Dholuo (Luo) speech corpus, pooled from two independent sources
and filtered to only genuinely transcribed audio. Part of the
AfroNet multi-language TTS data
effort.
Sources
Anv-ke/Dholuo — African Next
Voices, a pilot data-collection effort in Kenya led by the KenCorpus Consortium (a
coalition of Kenyan universities and research centers), funded by the Gates
Foundation. 91,672 clips, 186.1h, source = anv_ke. Gated on… See the full description on the dataset page: https://huggingface.co/datasets/Professor/dholuo-speech-data.english-dholuo_sentence-pairs_mt560
English-Dholuo Parallel Dataset
This dataset contains parallel sentences in English and Dholuo (Kenya).
Dataset Information
Language Pair: English ↔ Dholuo
Language Code: luo
Country: Kenya
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this dataset, please cite the citation guide… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-dholuo_sentence-pairs_mt560.Dholuo_POSDholuo_POS
