datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kinyarwanda_afrivoice_all_domains_v0.1
Afrivoice Kinyarwanda — All Domains
Combined dataset across 5 domains from the original source.
Note: the source dataset also includes a scripted_education domain, excluded
here due to a cluster of corrupted audio files in one of its shards.
Attribution
Original dataset: DigitalUmuganda/Afrivoice_Kinyarwanda
License: CC-BY-4.0
Attribution: Digital Umuganda
This dataset is derived from the above source and released under the same CC-BY-4.0 license.… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.safi-kinyarwanda-conversations
Safi Diction Kinyarwanda Conversational Speech Dataset
This dataset contains 1 hour of Kinyarwanda conversational speech collected using Safi's collection engine.
The recordings contain multiple speakers responding to survey questions. The original recordings were processed using speaker diarization to identify speaker turns. Consecutive turns from the same speaker were consolidated and split into speaker-specific audio clips of up to 15 seconds. These clips were then… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-kinyarwanda-conversations.kinyarwanda-speech-500h
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains 500 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~500 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-500h.kinyarwanda-speech-1000h
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains ~1000 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track B competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~1000 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-1000h.kinyarwanda-speech-sample
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains a sample from the 500 hours of Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~500 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-sample.KinyaWhisperDataset
Kinyarwanda Spoken Words Dataset
This dataset contains 102 short audio samples of spoken Kinyarwanda words, each labeled with its corresponding transcription. It is designed for training, evaluating, and experimenting with Automatic Speech Recognition (ASR) models in low-resource settings.
Structure
audio/: Contains 102 .wav files (mono, 16kHz)
transcripts.txt: Tab-separated transcription file (e.g., 001.wav\tmuraho)
manifest.jsonl: JSONL file with audio paths and text… See the full description on the dataset page: https://huggingface.co/datasets/benax-rw/KinyaWhisperDataset.kinyarwanda-speech-trimmed
Kinyarwanda Speech Dataset (Trimmed)
This dataset contains processed Kinyarwanda speech data with trimmed audio segments.
Dataset Structure
The dataset contains two splits:
dev_test: 9,263 samples
test: 9,265 samples
Features
Each sample contains:
id: Unique identifier for the sample
audio: Audio data
audio_language: Language of the audio (Kinyarwanda)
text: Transcription of the audio
prompt: Associated prompt or context
duration: Duration of the audio… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/kinyarwanda-speech-trimmed.TORGO-database
The TORGO Database: Acoustic and articulatory speech from speakers with dysarthria
Dataset Summary
This database only includes the short words and restricted sentence portion of the TORGO dataset.
For the full dataset which also includes non-words and unrestricted sentences please see: https://www.cs.toronto.edu/~complingweb/data/TORGO/torgo.html.
Transcripts have been normalized to remove punctuation but casing has been left. Few transcripts only had 'xxx' as text… See the full description on the dataset page: https://huggingface.co/datasets/kingp12/TORGO-database.kin-s-5
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains 500 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~500 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/kin-s-5.kinyarwanda-pastor-snac
Kinyarwanda Pastor Uwambaje — SNAC Dataset
Single-speaker Kinyarwanda TTS dataset from Pastor Uwambaje YouTube sermons.
Audio cleaned with htdemucs (music removal) + resemble-enhance (denoising).
Encoded with SNAC 24kHz, 7-token interleaved format.
Total uploaded: 13,380 clips at STOI >= 0.80 threshold (~18.4h).
STOI Quality Distribution
STOI Threshold
Clips
Hours
>= 0.80 (this dataset)
13,225
18.4h
>= 0.85
13,058
18.2h
>= 0.90
12,488
17.5h
>= 0.95
9,734… See the full description on the dataset page: https://huggingface.co/datasets/vysakh25/kinyarwanda-pastor-snac.Afrivoice_Kinyarwanda_old_version
Dataset summary
[need more information]
Supported tasks
[need more information]
How to use
[need more information]
Dataset structure
Data fields
[need more information]
Data splits
[need more information]
Data preprocessing
[need more information]
Licensing Information
All datasets are licensed under the Creative Commons license (CC-BY-4).
