CoolFace
Datasetpublic

speedykom-group/turkana-speech-dataset

Turkana Speech Dataset Speech dataset for Turkana (tuv) — Eastern Nilotic language, ~1M speakers, Kenya. Property Value Format WAV, 16 kHz, mono / UTF-8 transcripts Clips 5,151 segments Splits Train: 3,090 (60%) · Validation: 1,030 (20%) · Test: 1,031 (20%) — seed 42 Source GRN Bible narratives (Global Recordings Network, LLL series 1–8), segmented via silence detection Transcription Auto-generated via facebook/mms-1b-all (Teso adapter)… See the full description on the dataset page: https://huggingface.co/datasets/speedykom-group/turkana-speech-dataset.

sourceHugging Facecc-by-nc-4.0updated 2mo agoView on Hugging Face
1likes28downloads
Dataset Card

Turkana Speech Dataset

<img src="https://speedykom.de/speedykom-small.png" alt="Speedykom" width="150"/>

Speech dataset for Turkana (tuv) — Eastern Nilotic language, ~1M speakers, Kenya.
PropertyValue
FormatWAV, 16 kHz, mono / UTF-8 transcripts
Clips5,151 segments
SplitsTrain: 3,090 (60%) · Validation: 1,030 (20%) · Test: 1,031 (20%) — seed 42
SourceGRN Bible narratives (Global Recordings Network, LLL series 1–8), segmented via silence detection
TranscriptionAuto-generated via facebook/mms-1b-all (Teso adapter)

About

This dataset was created by Speedykom as part of an effort to advance speech technology for underserved African languages. Turkana is spoken by approximately one million people in northwestern Kenya and belongs to the Ateker (Teso-Turkana) language cluster within the Eastern Nilotic family.

Usage

python
from datasets import load_dataset

ds = load_dataset("speedykom-group/turkana-speech-dataset")
print(ds["train"][0])

Notes

Please respect the original license terms.

Citation

Speedykom - turkana-speech-dataset
https://huggingface.co/datasets/speedykom-group/turkana-speech-dataset
Created by Speedykom (https://speedykom.de)

Created by [Speedykom](https://speedykom.de)