CoolFace
Datasetpublic

badrex/kinyarwanda-speech-1000h

Kinyarwanda Automatic Speech Recognition Dataset Dataset Description This dataset contains ~1000 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track B competition. Dataset Details Language: Kinyarwanda (rw) Task: Automatic Speech Recognition Size: ~1000 hours of transcribed speech Domains: Health, Government, Financial Services… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-1000h.

sourceHugging Faceccupdated 1y agoView on Hugging Face
0likes211downloads
Dataset Card

Kinyarwanda Automatic Speech Recognition Dataset

Dataset Description

This dataset contains ~1000 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track B competition.

Dataset Details

  • Language: Kinyarwanda (rw)
  • Task: Automatic Speech Recognition
  • Size: ~1000 hours of transcribed speech
  • Domains: Health, Government, Financial Services, Education, Agriculture
  • Format: Audio files with corresponding transcriptions
  • Source: Created by Digital Umuganda with Gates Foundation funding

Dataset Structure

python
# example usage
from datasets import load_dataset

dataset = load_dataset("badrex/kinyarwanda-speech-500h")

Use Cases

  • training ASR models for Kinyarwanda
  • fine-tuning existing speech recognition models (e.g., Whisper)
  • research in low-resource speech recognition
  • building voice applications for Kinyarwanda speakers

License

The dataset is available under Creative Commons Attribution 4.0 (CC BY 4.0) license.

Citation

bibtex
@misc{kinyarwanda_asr_track_b,
  title={Kinyarwanda Automatic Speech Recognition Track B},
  author={Digital Umuganda},
  year={2025},
  url={https://www.kaggle.com/competitions/kinyarwanda-automatic-speech-recognition-track-b}
}