datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
afrispeak_kinyarwanda_male_tts_datasetkinyarwanda_afrivoice_all_domains_v0.1
Afrivoice Kinyarwanda — All Domains
Combined dataset across 5 domains from the original source.
Note: the source dataset also includes a scripted_education domain, excluded
here due to a cluster of corrupted audio files in one of its shards.
Attribution
Original dataset: DigitalUmuganda/Afrivoice_Kinyarwanda
License: CC-BY-4.0
Attribution: Digital Umuganda
This dataset is derived from the above source and released under the same CC-BY-4.0 license.… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.kinyarwanda_afrivoice_all_domains_v0.2
Kinyarwanda AfriVoice — All Domains (v0.2)
Cleaned version of ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.
Changes from v0.1
Removed rows with empty/null transcription values across all splits (train/validation/test)
Audio and domain labels unchanged; only null-transcription rows were dropped
Source
Original data from DigitalUmuganda/Afrivoice_Kinyarwanda (CC-BY-4.0),
extracted and concatenated across 5 domains (agriculture, education… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.2.safi-kinyarwanda-conversations
Safi Diction Kinyarwanda Conversational Speech Dataset
This dataset contains 1 hour of Kinyarwanda conversational speech collected using Safi's collection engine.
The recordings contain multiple speakers responding to survey questions. The original recordings were processed using speaker diarization to identify speaker turns. Consecutive turns from the same speaker were consolidated and split into speaker-specific audio clips of up to 15 seconds. These clips were then… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-kinyarwanda-conversations.kinyarwanda-tts-dataset
Kinyarwanda TTS dataset
The dataset consists of 3992 clips of Kinyarwanda TTS corpus recorded in a studio using a voice actress, it was collected in the mbaza project
Data structure
Audio: 3992 Single voice studio recordings by a voice actress
Text: CSV with audio name and corresponding written text
Language
The corresponding dataset is in the Kinyarwanda Language
Dataset Creation
Text collected had to include Kinyarwanda syllabes, which is made by… See the full description on the dataset page: https://huggingface.co/datasets/mbazaNLP/kinyarwanda-tts-dataset.kinyarwanda-speech-500h
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains 500 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~500 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-500h.kinyarwanda-asr-track-akinyarwanda-speech-1000h
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains ~1000 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track B competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~1000 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-1000h.kinyarwanda-tts-dataset-kin
Kinyarwanda TTS Dataset (Split Version)
This dataset is a reformatted version of the mbazaNLP/kinyarwanda-tts-dataset.
Modifications
Original data was provided as a single set of 3,992 clips.
This version has been split into Train (80%), Validation (10%), and Test (10%) sets.
Audio files have been processed into the Hugging Face datasets format for easier loading.
Credits & Acknowledgements
Original data created and provided by Mbaza NLP. All credit for the… See the full description on the dataset page: https://huggingface.co/datasets/Professor/kinyarwanda-tts-dataset-kin.Afrivoice_Kinyarwanda_ASR_clonekinyarwanda_denoisedafrispeak_kinyarwanda_female_tts_datasetkinyarwanda_cleaned_testset_verified_20HRSkinyarwanda_cleaned_testset_verified_200HRSrwandan_kinyarwanda_nonstandard_speech_v1.0This dataset provides 61.7 hours of Kinyarwanda speech recordings (14,739 samples) from 61 Rwandan speakers living with speech impairments. The participants represent a limited diversity of speech patterns, mostly stuttering, and a few examples of Dysarthria, Dysphonia, and Phonological disorders.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or phrase level. All speech recordings of this datasets have been… See the full description on the dataset page: https://huggingface.co/datasets/cdli/rwandan_kinyarwanda_nonstandard_speech_v1.0.Afrivoice-Kinyarwanda-ASRkinyarwanda-speech-sample
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains a sample from the 500 hours of Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~500 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-sample.kinyarwanda-mms-datakinyarwanda_cleaned_testset_verifiedkinyarwanda-male-youtube2-snackinyarwanda_cleaned_testset_verified_100HRSigisha-kinyarwanda-asr
language:
- rw
- en
license: cc-by-4.0
task_categories:
- automatic-speech-recognition
- translation
task_ids:
- speech-recognition
- speech-translation
size_categories:
- 1K<n<10K
tags:
- kinyarwanda
- rwanda
- speech
- whisper
- igisha
- low-resource
pretty_name: Igisha Kinyarwanda ASR & Translation
Igisha Kinyarwanda Speech Dataset
A curated Kinyarwanda speech dataset collected for fine-tuning
Whisper on transcription
(Kinyarwanda → Kinyarwanda text) and… See the full description on the dataset page: https://huggingface.co/datasets/LeonceNsh/igisha-kinyarwanda-asr.kinyarwanda-speech-trimmed
Kinyarwanda Speech Dataset (Trimmed)
This dataset contains processed Kinyarwanda speech data with trimmed audio segments.
Dataset Structure
The dataset contains two splits:
dev_test: 9,263 samples
test: 9,265 samples
Features
Each sample contains:
id: Unique identifier for the sample
audio: Audio data
audio_language: Language of the audio (Kinyarwanda)
text: Transcription of the audio
prompt: Associated prompt or context
duration: Duration of the audio… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/kinyarwanda-speech-trimmed.kinyarwanda-mary-snackinyarwanda-tts-splitkinyarwanda_cleaned_testset_verified_10HRSkinyarwanda-speech-hackathon
📚 Kinyarwanda ASR Dataset
This dataset contains transcribed Kinyarwanda audio, designed to support training and evaluation of Automatic Speech Recognition (ASR) systems. It is part of a study on how varying training data volumes affect model performance using Whisper-large-v3.
📂 Data Overview
The full dataset consists of approximately 263,000 audio samples covering 5 key domains:
🏥 Health
🏛️ Government
💰 Financial Services
🎓 Education
🌾 Agriculture
To… See the full description on the dataset page: https://huggingface.co/datasets/evie-8/kinyarwanda-speech-hackathon.youtube-kinyarwanda-snac-scoredkinyarwanda_datasetkinyarwanda-hackathon
