Tonykip/kenyan-swahili-asr-clean
Kenyan Swahili ASR — clean, Nemotron-ready A cleaned, validated Kenyan-Swahili ASR corpus prepared for fine-tuning streaming ASR models (e.g. NVIDIA Nemotron 3.5 ASR). This is an initial ~30h subset for pipeline validation; a larger version will follow. Format Audio: 16 kHz mono WAV (embedded). text: cased + punctuated transcript. target_lang: sw-KE · source, dialect, duration columns included. Source & license Derived from Afrivoice… See the full description on the dataset page: https://huggingface.co/datasets/Tonykip/kenyan-swahili-asr-clean.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face