datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KinyaWhisperDataset
Kinyarwanda Spoken Words Dataset
This dataset contains 102 short audio samples of spoken Kinyarwanda words, each labeled with its corresponding transcription. It is designed for training, evaluating, and experimenting with Automatic Speech Recognition (ASR) models in low-resource settings.
Structure
audio/: Contains 102 .wav files (mono, 16kHz)
transcripts.txt: Tab-separated transcription file (e.g., 001.wav\tmuraho)
manifest.jsonl: JSONL file with audio paths and text… See the full description on the dataset page: https://huggingface.co/datasets/benax-rw/KinyaWhisperDataset.rw-tts-dataset
Rw Tts Dataset
Dataset Description
Kinyarwanda (rw) text-to-speech dataset. Studio-recorded read speech aligned with transcriptions, collected in Rwanda.
Languages
Language: Kinyarwanda (rw)
BCP-47: rw
Source tag
rw — identifies the origin of each sample in the source column.
Dataset Structure
Column
Type
Description
audio
Audio
Raw WAV audio at original recording frequency
text
string
Transcription of the spoken content… See the full description on the dataset page: https://huggingface.co/datasets/Bateesa/rw-tts-dataset.
