CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DigitalUmuganda /afrispeak_kinyarwanda_male_tts_datasetgatedaudio0 likes531 downloads2y agoHugging Face02ElizabethMwangi /kinyarwanda_afrivoice_all_domains_v0.1 Afrivoice Kinyarwanda — All Domains Combined dataset across 5 domains from the original source. Note: the source dataset also includes a scripted_education domain, excluded here due to a cluster of corrupted audio files in one of its shards. Attribution Original dataset: DigitalUmuganda/Afrivoice_Kinyarwanda License: CC-BY-4.0 Attribution: Digital Umuganda This dataset is derived from the above source and released under the same CC-BY-4.0 license.… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.audioautomatic-speech-recognition100K<n<1M0 likes520 downloads2mo agoHugging Face03ElizabethMwangi /kinyarwanda_afrivoice_all_domains_v0.2 Kinyarwanda AfriVoice — All Domains (v0.2) Cleaned version of ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1. Changes from v0.1 Removed rows with empty/null transcription values across all splits (train/validation/test) Audio and domain labels unchanged; only null-transcription rows were dropped Source Original data from DigitalUmuganda/Afrivoice_Kinyarwanda (CC-BY-4.0), extracted and concatenated across 5 domains (agriculture, education… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.2.audio100K<n<1M0 likes388 downloads1mo agoHugging Face04martinturuta /safi-kinyarwanda-conversations Safi Diction Kinyarwanda Conversational Speech Dataset This dataset contains 1 hour of Kinyarwanda conversational speech collected using Safi's collection engine. The recordings contain multiple speakers responding to survey questions. The original recordings were processed using speaker diarization to identify speaker turns. Consecutive turns from the same speaker were consolidated and split into speaker-specific audio clips of up to 15 seconds. These clips were then… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-kinyarwanda-conversations.audioautomatic-speech-recognitionn<1K0 likes339 downloads20d agoHugging Face05mbazaNLP /kinyarwanda-tts-datasetgated Kinyarwanda TTS dataset The dataset consists of 3992 clips of Kinyarwanda TTS corpus recorded in a studio using a voice actress, it was collected in the mbaza project Data structure Audio: 3992 Single voice studio recordings by a voice actress Text: CSV with audio name and corresponding written text Language The corresponding dataset is in the Kinyarwanda Language Dataset Creation Text collected had to include Kinyarwanda syllabes, which is made by… See the full description on the dataset page: https://huggingface.co/datasets/mbazaNLP/kinyarwanda-tts-dataset.audio1K<n<10K5 likes301 downloads3y agoHugging Face06badrex /kinyarwanda-speech-500h Kinyarwanda Automatic Speech Recognition Dataset Dataset Description This dataset contains 500 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition. Dataset Details Language: Kinyarwanda (rw) Task: Automatic Speech Recognition Size: ~500 hours of transcribed speech Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-500h.audioautomatic-speech-recognition10K<n<100K0 likes300 downloads1y agoHugging Face07Oluwadara /kinyarwanda-asr-track-aaudio10K<n<100K0 likes252 downloads1y agoHugging Face08badrex /kinyarwanda-speech-1000h Kinyarwanda Automatic Speech Recognition Dataset Dataset Description This dataset contains ~1000 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track B competition. Dataset Details Language: Kinyarwanda (rw) Task: Automatic Speech Recognition Size: ~1000 hours of transcribed speech Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-1000h.audioautomatic-speech-recognition100K<n<1M0 likes211 downloads1y agoHugging Face09Professor /kinyarwanda-tts-dataset-kin Kinyarwanda TTS Dataset (Split Version) This dataset is a reformatted version of the mbazaNLP/kinyarwanda-tts-dataset. Modifications Original data was provided as a single set of 3,992 clips. This version has been split into Train (80%), Validation (10%), and Test (10%) sets. Audio files have been processed into the Hugging Face datasets format for easier loading. Credits & Acknowledgements Original data created and provided by Mbaza NLP. All credit for the… See the full description on the dataset page: https://huggingface.co/datasets/Professor/kinyarwanda-tts-dataset-kin.audiotext-to-speech1K<n<10K0 likes118 downloads8mo agoHugging Face10ThatDev /Afrivoice_Kinyarwanda_ASR_cloneaudio100K<n<1M0 likes96 downloads5mo agoHugging Face11tianyiwordly /kinyarwanda_denoisedaudio100K<n<1M0 likes93 downloads2y agoHugging Face12DigitalUmuganda /afrispeak_kinyarwanda_female_tts_datasetaudio0 likes84 downloads2y agoHugging Face13KYAGABA /kinyarwanda_cleaned_testset_verified_20HRSaudio10K<n<100K0 likes50 downloads2y agoHugging Face14KYAGABA /kinyarwanda_cleaned_testset_verified_200HRSaudio100K<n<1M0 likes40 downloads2y agoHugging Face15cdli /rwandan_kinyarwanda_nonstandard_speech_v1.0gatedThis dataset provides 61.7 hours of Kinyarwanda speech recordings (14,739 samples) from 61 Rwandan speakers living with speech impairments. The participants represent a limited diversity of speech patterns, mostly stuttering, and a few examples of Dysarthria, Dysphonia, and Phonological disorders. This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or phrase level. All speech recordings of this datasets have been… See the full description on the dataset page: https://huggingface.co/datasets/cdli/rwandan_kinyarwanda_nonstandard_speech_v1.0.audio10K<n<100K0 likes38 downloads1mo agoHugging Face16Kira-Floris /Afrivoice-Kinyarwanda-ASRaudio100K<n<1M0 likes31 downloads6mo agoHugging Face17badrex /kinyarwanda-speech-sample Kinyarwanda Automatic Speech Recognition Dataset Dataset Description This dataset contains a sample from the 500 hours of Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition. Dataset Details Language: Kinyarwanda (rw) Task: Automatic Speech Recognition Size: ~500 hours of transcribed speech Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-sample.audioautomatic-speech-recognition1K<n<10K0 likes24 downloads1y agoHugging Face18Professor /kinyarwanda-mms-dataaudio10K<n<100K0 likes23 downloads8mo agoHugging Face19KYAGABA /kinyarwanda_cleaned_testset_verifiedaudio100K<n<1M0 likes20 downloads2y agoHugging Face20vysakh25 /kinyarwanda-male-youtube2-snacaudion<1K0 likes18 downloads7mo agoHugging Face21KYAGABA /kinyarwanda_cleaned_testset_verified_100HRSaudio10K<n<100K0 likes17 downloads2y agoHugging Face22LeonceNsh /igisha-kinyarwanda-asr language: - rw - en license: cc-by-4.0 task_categories: - automatic-speech-recognition - translation task_ids: - speech-recognition - speech-translation size_categories: - 1K<n<10K tags: - kinyarwanda - rwanda - speech - whisper - igisha - low-resource pretty_name: Igisha Kinyarwanda ASR & Translation Igisha Kinyarwanda Speech Dataset A curated Kinyarwanda speech dataset collected for fine-tuning Whisper on transcription (Kinyarwanda → Kinyarwanda text) and… See the full description on the dataset page: https://huggingface.co/datasets/LeonceNsh/igisha-kinyarwanda-asr.audio1K<n<10K0 likes17 downloads9d agoHugging Face23yigagilbert /kinyarwanda-speech-trimmedgated Kinyarwanda Speech Dataset (Trimmed) This dataset contains processed Kinyarwanda speech data with trimmed audio segments. Dataset Structure The dataset contains two splits: dev_test: 9,263 samples test: 9,265 samples Features Each sample contains: id: Unique identifier for the sample audio: Audio data audio_language: Language of the audio (Kinyarwanda) text: Transcription of the audio prompt: Associated prompt or context duration: Duration of the audio… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/kinyarwanda-speech-trimmed.audioautomatic-speech-recognition10K<n<100K0 likes15 downloads1y agoHugging Face24vysakh25 /kinyarwanda-mary-snacaudio1K<n<10K0 likes12 downloads7mo agoHugging Face25luciayen /kinyarwanda-tts-splitaudio1K<n<10K0 likes11 downloads9mo agoHugging Face26KYAGABA /kinyarwanda_cleaned_testset_verified_10HRSaudio1K<n<10K0 likes8 downloads2y agoHugging Face27evie-8 /kinyarwanda-speech-hackathongated 📚 Kinyarwanda ASR Dataset This dataset contains transcribed Kinyarwanda audio, designed to support training and evaluation of Automatic Speech Recognition (ASR) systems. It is part of a study on how varying training data volumes affect model performance using Whisper-large-v3. 📂 Data Overview The full dataset consists of approximately 263,000 audio samples covering 5 key domains: 🏥 Health 🏛️ Government 💰 Financial Services 🎓 Education 🌾 Agriculture To… See the full description on the dataset page: https://huggingface.co/datasets/evie-8/kinyarwanda-speech-hackathon.audio100K<n<1M0 likes7 downloads1y agoHugging Face28vysakh25 /youtube-kinyarwanda-snac-scoredaudio1K<n<10K0 likes7 downloads7mo agoHugging Face29yiki-ui /kinyarwanda_datasetaudio10K<n<100K0 likes6 downloads1y agoHugging Face30evie-8 /kinyarwanda-hackathongatedaudio100K<n<1M0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.