datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Kenyan-Swahili-Speech
Read Speech in Kenyan Swahili (6h)
A single-speaker read speech dataset in Kenyan Swahili, containing approximately 6 hours of prompted recordings from an anonymous male speaker. The dataset was produced as part of CLEAR Global's Gamayun Language Data Kits initiative, which develops open-source language resources for under-resourced languages used in humanitarian contexts.
The sentence set is shared with CLEAR Global's Gamayun Swahili–English parallel text kit. English source… See the full description on the dataset page: https://huggingface.co/datasets/CLEAR-Global/Kenyan-Swahili-Speech.kenyan-swahili-asr-clean
Kenyan Swahili ASR — clean, Nemotron-ready
A cleaned, validated Kenyan-Swahili ASR corpus prepared for fine-tuning streaming ASR
models (e.g. NVIDIA Nemotron 3.5 ASR). This is an initial ~30h subset for
pipeline validation; a larger version will follow.
Format
Audio: 16 kHz mono WAV (embedded).
text: cased + punctuated transcript.
target_lang: sw-KE · source, dialect, duration columns included.
Source & license
Derived from Afrivoice… See the full description on the dataset page: https://huggingface.co/datasets/Tonykip/kenyan-swahili-asr-clean.cv17_sw_kenyan_sample
Common Voice 17.0 — Swahili (Kenyan Sample)
This dataset is a filtered sample of the Mozilla Common Voice 17.0 corpus, focusing on Swahili (sw) speech with Kenyan voices.
It has been subsetted for experimentation and prototyping in ASR (Automatic Speech Recognition) models targeting speech-impaired users in Kenya, covering Kenyan English and Kiswahili.
Dataset Summary
Language: Kiswahili (Swahili, sw)
Accent/Region: Kenyan speakers
Domain: Conversational… See the full description on the dataset page: https://huggingface.co/datasets/Veronica1NW/cv17_sw_kenyan_sample.Kenyan-Swahili-Speech
Read Speech in Kenyan Swahili (6h)
A single-speaker read speech dataset in Kenyan Swahili, containing approximately 6 hours of prompted recordings from an anonymous male speaker. The dataset was produced as part of CLEAR Global's Gamayun Language Data Kits initiative, which develops open-source language resources for under-resourced languages used in humanitarian contexts.
The sentence set is shared with CLEAR Global's Gamayun Swahili–English parallel text kit. English source… See the full description on the dataset page: https://huggingface.co/datasets/Nzyoka19/Kenyan-Swahili-Speech.
