Tonykip/kenyan-swahili-asr-clean
Kenyan Swahili ASR — clean, Nemotron-ready A cleaned, validated Kenyan-Swahili ASR corpus prepared for fine-tuning streaming ASR models (e.g. NVIDIA Nemotron 3.5 ASR). This is an initial ~30h subset for pipeline validation; a larger version will follow. Format Audio: 16 kHz mono WAV (embedded). text: cased + punctuated transcript. target_lang: sw-KE · source, dialect, duration columns included. Source & license Derived from Afrivoice… See the full description on the dataset page: https://huggingface.co/datasets/Tonykip/kenyan-swahili-asr-clean.
Kenyan Swahili ASR — clean, Nemotron-ready
A cleaned, validated Kenyan-Swahili ASR corpus prepared for fine-tuning streaming ASR models (e.g. NVIDIA Nemotron 3.5 ASR). This is an initial ~30h subset for pipeline validation; a larger version will follow.
Format
- Audio: 16 kHz mono WAV (embedded).
text: cased + punctuated transcript.target_lang:sw-KE·source,dialect,durationcolumns included.
Source & license
- Derived from Afrivoice (DigitalUmuganda) Swahili, via the
badrex/swahili-speech-400hrclean variant. Original license CC-BY-4.0; this derivative is released under CC-BY-4.0 with attribution to DigitalUmuganda/Afrivoice and badrex.
Cleaning (this release)
Resampled to 16k mono; dropped clips <0.5s / >30s, empty transcripts, and implausible speech-rate (>4 words/s); exact-transcript dedup. Kept 5106 clips ≈ 30.00h. Drop counts: {'short': 0, 'long': 85, 'empty': 0, 'rate': 0, 'dup': 0}.
Validation (pre-clean, n=50 sample)
Independent checks on the source: 100% language-ID Swahili; Whisper-large-v3 cross-check median CER ≈ 0.075 (transcript matches audio); 100% cased + punctuated.
Intended use
Fine-tuning / evaluating Swahili ASR. Not for the identification of individuals.
