CoolFace
Datasetpublic

Tonykip/kenyan-swahili-asr-clean

Kenyan Swahili ASR — clean, Nemotron-ready A cleaned, validated Kenyan-Swahili ASR corpus prepared for fine-tuning streaming ASR models (e.g. NVIDIA Nemotron 3.5 ASR). This is an initial ~30h subset for pipeline validation; a larger version will follow. Format Audio: 16 kHz mono WAV (embedded). text: cased + punctuated transcript. target_lang: sw-KE · source, dialect, duration columns included. Source & license Derived from Afrivoice… See the full description on the dataset page: https://huggingface.co/datasets/Tonykip/kenyan-swahili-asr-clean.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes28downloads
Dataset Card

Kenyan Swahili ASR — clean, Nemotron-ready

A cleaned, validated Kenyan-Swahili ASR corpus prepared for fine-tuning streaming ASR models (e.g. NVIDIA Nemotron 3.5 ASR). This is an initial ~30h subset for pipeline validation; a larger version will follow.

Format

  • —Audio: 16 kHz mono WAV (embedded).
  • —text: cased + punctuated transcript.
  • —target_lang: sw-KE · source, dialect, duration columns included.

Source & license

  • —Derived from Afrivoice (DigitalUmuganda) Swahili, via the badrex/swahili-speech-400hr clean variant. Original license CC-BY-4.0; this derivative is released under CC-BY-4.0 with attribution to DigitalUmuganda/Afrivoice and badrex.

Cleaning (this release)

Resampled to 16k mono; dropped clips <0.5s / >30s, empty transcripts, and implausible speech-rate (>4 words/s); exact-transcript dedup. Kept 5106 clips ≈ 30.00h. Drop counts: {'short': 0, 'long': 85, 'empty': 0, 'rate': 0, 'dup': 0}.

Validation (pre-clean, n=50 sample)

Independent checks on the source: 100% language-ID Swahili; Whisper-large-v3 cross-check median CER ≈ 0.075 (transcript matches audio); 100% cased + punctuated.

Intended use

Fine-tuning / evaluating Swahili ASR. Not for the identification of individuals.