datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
afrispeak_kinyarwanda_male_tts_datasetkinyarwanda_afrivoice_all_domains_v0.1
Afrivoice Kinyarwanda — All Domains
Combined dataset across 5 domains from the original source.
Note: the source dataset also includes a scripted_education domain, excluded
here due to a cluster of corrupted audio files in one of its shards.
Attribution
Original dataset: DigitalUmuganda/Afrivoice_Kinyarwanda
License: CC-BY-4.0
Attribution: Digital Umuganda
This dataset is derived from the above source and released under the same CC-BY-4.0 license.… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.kinyarwanda_afrivoice_all_domains_v0.2
Kinyarwanda AfriVoice — All Domains (v0.2)
Cleaned version of ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.
Changes from v0.1
Removed rows with empty/null transcription values across all splits (train/validation/test)
Audio and domain labels unchanged; only null-transcription rows were dropped
Source
Original data from DigitalUmuganda/Afrivoice_Kinyarwanda (CC-BY-4.0),
extracted and concatenated across 5 domains (agriculture, education… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.2.kinetics-700-2020common_voice_1000_KINkinetics-400safi-kinyarwanda-conversations
Safi Diction Kinyarwanda Conversational Speech Dataset
This dataset contains 1 hour of Kinyarwanda conversational speech collected using Safi's collection engine.
The recordings contain multiple speakers responding to survey questions. The original recordings were processed using speaker diarization to identify speaker turns. Consecutive turns from the same speaker were consolidated and split into speaker-specific audio clips of up to 15 seconds. These clips were then… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-kinyarwanda-conversations.Honor-of-Kings-Audio该数据集包含王者荣耀中所有英雄及其所有皮肤的语音数据。数据组织结构如下:
每个顶层文件夹以英雄名称命名;
英雄文件夹下的子文件夹为该英雄的各个皮肤名称;
每个皮肤文件夹内包含若干语音音频文件。
This dataset contains voice recordings for all heroes and all skins from Honor of Kings. The data is organized as follows:
Each top-level folder is named after a hero;
Subfolders within a hero folder correspond to that hero’s skins;
Each skin folder contains multiple audio files of voice lines.
kinyarwanda-speech-500h
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains 500 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~500 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-500h.kinya-ag-tts
Kinyarwanda Agricultural Text-to-Speech Dataset
In Rwanda, many farmers struggle to access timely, personalized agricultural information. Traditional channels - like radio, TV, and online sources - offer limited reach and interactivity, while extension services and a national call center, staffed by only two agents for over two million farmers, face capacity constraints. To address these gaps, we developed a 24/7 AI-enabled Interactive Voice Response (IVR) tool. Accessible via a… See the full description on the dataset page: https://huggingface.co/datasets/C4IR-RW/kinya-ag-tts.kinyarwanda-asr-track-akinyarwanda-speech-1000h
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains ~1000 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track B competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~1000 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-1000h.kinyarwanda-tts-dataset
Kinyarwanda TTS dataset
The dataset consists of 3992 clips of Kinyarwanda TTS corpus recorded in a studio using a voice actress, it was collected in the mbaza project
Data structure
Audio: 3992 Single voice studio recordings by a voice actress
Text: CSV with audio name and corresponding written text
Language
The corresponding dataset is in the Kinyarwanda Language
Dataset Creation
Text collected had to include Kinyarwanda syllabes, which is made by… See the full description on the dataset page: https://huggingface.co/datasets/mbazaNLP/kinyarwanda-tts-dataset.kin_cleaned_commonvoiceAfrivoice_Kinyarwanda_ASR_clonekinyarwanda-tts-dataset-kin
Kinyarwanda TTS Dataset (Split Version)
This dataset is a reformatted version of the mbazaNLP/kinyarwanda-tts-dataset.
Modifications
Original data was provided as a single set of 3,992 clips.
This version has been split into Train (80%), Validation (10%), and Test (10%) sets.
Audio files have been processed into the Hugging Face datasets format for easier loading.
Credits & Acknowledgements
Original data created and provided by Mbaza NLP. All credit for the… See the full description on the dataset page: https://huggingface.co/datasets/Professor/kinyarwanda-tts-dataset-kin.kinh-phap-hoa-ke-trom-huongNormalized using https://github.com/oysterlanguage/emiliapipex
@article{emilia,
title={Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation},
author={He, Haorui and Shang, Zengqiang and Wang, Chaoren and Li, Xuyuan and Gu, Yicheng and Hua, Hua and Liu, Liwei and Yang, Chen and Li, Jiaqi and Shi, Peiyang and Wang, Yuancheng and Chen, Kai and Zhang, Pengyuan and Wu, Zhizheng},
journal={arXiv},
volume={abs/2407.05361}… See the full description on the dataset page: https://huggingface.co/datasets/hr16/kinh-phap-hoa-ke-trom-huong.afrispeak_kinyarwanda_female_tts_datasetkinyarwanda_denoisedkinetics-600kinyarwanda_cleaned_testset_verified_20HRSkin_cleaned_commonvoice_rwanda_200hoursEnglish-United-Kingdom-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 90,334 hours of processed English (UK) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, interruptions, and natural speaking behavior commonly… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English-United-Kingdom-Call-Center-Audio-Dataset-Single-Channel.rwandan_kinyarwanda_nonstandard_speech_v1.0This dataset provides 61.7 hours of Kinyarwanda speech recordings (14,739 samples) from 61 Rwandan speakers living with speech impairments. The participants represent a limited diversity of speech patterns, mostly stuttering, and a few examples of Dysarthria, Dysphonia, and Phonological disorders.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or phrase level. All speech recordings of this datasets have been… See the full description on the dataset page: https://huggingface.co/datasets/cdli/rwandan_kinyarwanda_nonstandard_speech_v1.0.English_United_Kingdom_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 90,334 hours of processed English (UK) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English_United_Kingdom_Call_Center_Audio_Dataset_Dual_Channel.kinyarwanda_cleaned_testset_verified_200HRSAfrivoice-Kinyarwanda-ASRkinyarwanda-speech-sample
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains a sample from the 500 hours of Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~500 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-sample.Red-Alert-2-Full-Voice-DataExtract from Red Alert 2 & Yuri's Revenge.
Include characters:
General Carville
President Dugan
Professor Einstein
Yuri battlefield controller
Eva
Premier Romanov
Agent Tanya
Yuri
Zofia
2191 WAV files in total.
PCM_f32le 16bit 22050hz
Ready for VITS-Fast-Fine-Tuning training.
jambal_common_voice
