datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kinyarwanda_afrivoice_all_domains_v0.1
Afrivoice Kinyarwanda — All Domains
Combined dataset across 5 domains from the original source.
Note: the source dataset also includes a scripted_education domain, excluded
here due to a cluster of corrupted audio files in one of its shards.
Attribution
Original dataset: DigitalUmuganda/Afrivoice_Kinyarwanda
License: CC-BY-4.0
Attribution: Digital Umuganda
This dataset is derived from the above source and released under the same CC-BY-4.0 license.… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.safi-kinyarwanda-conversations
Safi Diction Kinyarwanda Conversational Speech Dataset
This dataset contains 1 hour of Kinyarwanda conversational speech collected using Safi's collection engine.
The recordings contain multiple speakers responding to survey questions. The original recordings were processed using speaker diarization to identify speaker turns. Consecutive turns from the same speaker were consolidated and split into speaker-specific audio clips of up to 15 seconds. These clips were then… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-kinyarwanda-conversations.kinyarwanda-speech-500h
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains 500 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~500 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-500h.kinyarwanda-speech-1000h
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains ~1000 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track B competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~1000 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-1000h.Afrivoice_Kinyarwanda
Dataset Card for the image text and voice dataset
Dataset Description
Each datapoint in this dataset consists of a JPEG image, a corresponding audio Webm file describing the image, and when available, the transcription of the audio file.
Domain
Total Hours
Transcribed Hours
Number of Clips
Dataset Size (GB)
Agriculture
467.13
465.40
86,305
30.13
Health
994.32
992.87
179,219
58.53
Finance
564.21
563.11
103,159
38.55
Government
676.10
674.11
122,265
49.22… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/Afrivoice_Kinyarwanda.fleurs-kinyarwanda
Fleur Kinyarwanda dataset
Fleur is a multilingual text and audio dataset. The original dataset was created by Google . The dataset can be used when building speech to text, speech to text translation and speech to speech translation. It is a good tool to benchmark speech application especially across languages. As of present Kinyarwanda did not have a fleur dataset hindering opportunities for building Kinyarwanda speech technology.
This dataset was created by 29 linguists that… See the full description on the dataset page: https://huggingface.co/datasets/mbazaNLP/fleurs-kinyarwanda.kinyarwanda-speech-sample
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains a sample from the 500 hours of Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~500 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-sample.kinyarwanda-speech-trimmed
Kinyarwanda Speech Dataset (Trimmed)
This dataset contains processed Kinyarwanda speech data with trimmed audio segments.
Dataset Structure
The dataset contains two splits:
dev_test: 9,263 samples
test: 9,265 samples
Features
Each sample contains:
id: Unique identifier for the sample
audio: Audio data
audio_language: Language of the audio (Kinyarwanda)
text: Transcription of the audio
prompt: Associated prompt or context
duration: Duration of the audio… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/kinyarwanda-speech-trimmed.kinyarwanda-speech-data
Kinyarwanda Speech Data (Pooled)
A ~989.9-hour pooled Kinyarwanda speech corpus, combining two independently-sourced
datasets into one consistently-formatted corpus for speech modeling (TTS / ASR). Part
of the AfroNet multi-language TTS data
effort — sibling release to the Yoruba/Hausa/Igbo pools, but sourced entirely
differently: DSN African Voices, NaijaVoices, and WAXAL (the sources behind the other
three languages) don't cover Kinyarwanda at all.
Sources… See the full description on the dataset page: https://huggingface.co/datasets/Professor/kinyarwanda-speech-data.kinyarwanda-pastor-snac
Kinyarwanda Pastor Uwambaje — SNAC Dataset
Single-speaker Kinyarwanda TTS dataset from Pastor Uwambaje YouTube sermons.
Audio cleaned with htdemucs (music removal) + resemble-enhance (denoising).
Encoded with SNAC 24kHz, 7-token interleaved format.
Total uploaded: 13,380 clips at STOI >= 0.80 threshold (~18.4h).
STOI Quality Distribution
STOI Threshold
Clips
Hours
>= 0.80 (this dataset)
13,225
18.4h
>= 0.85
13,058
18.2h
>= 0.90
12,488
17.5h
>= 0.95
9,734… See the full description on the dataset page: https://huggingface.co/datasets/vysakh25/kinyarwanda-pastor-snac.Afrivoice_Kinyarwanda_old_version
Dataset summary
[need more information]
Supported tasks
[need more information]
How to use
[need more information]
Dataset structure
Data fields
[need more information]
Data splits
[need more information]
Data preprocessing
[need more information]
Licensing Information
All datasets are licensed under the Creative Commons license (CC-BY-4).
