speedykom-group/turkana-speech-dataset
Turkana Speech Dataset Speech dataset for Turkana (tuv) — Eastern Nilotic language, ~1M speakers, Kenya. Property Value Format WAV, 16 kHz, mono / UTF-8 transcripts Clips 5,151 segments Splits Train: 3,090 (60%) · Validation: 1,030 (20%) · Test: 1,031 (20%) — seed 42 Source GRN Bible narratives (Global Recordings Network, LLL series 1–8), segmented via silence detection Transcription Auto-generated via facebook/mms-1b-all (Teso adapter)… See the full description on the dataset page: https://huggingface.co/datasets/speedykom-group/turkana-speech-dataset.
Turkana Speech Dataset
<img src="https://speedykom.de/speedykom-small.png" alt="Speedykom" width="150"/>
Speech dataset for Turkana (tuv) — Eastern Nilotic language, ~1M speakers, Kenya.
About
This dataset was created by Speedykom as part of an effort to advance speech technology for underserved African languages. Turkana is spoken by approximately one million people in northwestern Kenya and belongs to the Ateker (Teso-Turkana) language cluster within the Eastern Nilotic family.
Usage
from datasets import load_dataset
ds = load_dataset("speedykom-group/turkana-speech-dataset")
print(ds["train"][0])Notes
- Transcriptions were generated using the Teso (teo) ASR adapter from
facebook/mms-1b-all— the closest available language to Turkana. Manual review and correction is recommended. - Audio sourced from publicly available GRN recordings (LLL series 1–8):
- LLL 1 — Beginning with GOD
- LLL 2 — Mighty Men of GOD
- LLL 3 — Victory through GOD
- LLL 4 — Servants of GOD
- LLL 5 — On Trial for GOD
- LLL 6 — JESUS Teacher Healer
- LLL 7 — JESUS Lord Saviour
- LLL 8 — Acts of the HOLY SPIRIT
Please respect the original license terms.
Citation
Speedykom - turkana-speech-dataset
https://huggingface.co/datasets/speedykom-group/turkana-speech-dataset
Created by Speedykom (https://speedykom.de)Created by [Speedykom](https://speedykom.de)
