CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ElizabethMwangi /kinyarwanda_afrivoice_all_domains_v0.1 Afrivoice Kinyarwanda — All Domains Combined dataset across 5 domains from the original source. Note: the source dataset also includes a scripted_education domain, excluded here due to a cluster of corrupted audio files in one of its shards. Attribution Original dataset: DigitalUmuganda/Afrivoice_Kinyarwanda License: CC-BY-4.0 Attribution: Digital Umuganda This dataset is derived from the above source and released under the same CC-BY-4.0 license.… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.audioautomatic-speech-recognition100K<n<1M0 likes461 downloads2mo agoHugging Face02martinturuta /safi-kinyarwanda-conversations Safi Diction Kinyarwanda Conversational Speech Dataset This dataset contains 1 hour of Kinyarwanda conversational speech collected using Safi's collection engine. The recordings contain multiple speakers responding to survey questions. The original recordings were processed using speaker diarization to identify speaker turns. Consecutive turns from the same speaker were consolidated and split into speaker-specific audio clips of up to 15 seconds. These clips were then… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-kinyarwanda-conversations.audioautomatic-speech-recognitionn<1K0 likes347 downloads22d agoHugging Face03badrex /kinyarwanda-speech-500h Kinyarwanda Automatic Speech Recognition Dataset Dataset Description This dataset contains 500 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition. Dataset Details Language: Kinyarwanda (rw) Task: Automatic Speech Recognition Size: ~500 hours of transcribed speech Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-500h.audioautomatic-speech-recognition10K<n<100K0 likes313 downloads1y agoHugging Face04badrex /kinyarwanda-speech-1000h Kinyarwanda Automatic Speech Recognition Dataset Dataset Description This dataset contains ~1000 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track B competition. Dataset Details Language: Kinyarwanda (rw) Task: Automatic Speech Recognition Size: ~1000 hours of transcribed speech Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-1000h.audioautomatic-speech-recognition100K<n<1M0 likes217 downloads1y agoHugging Face05badrex /kinyarwanda-speech-sample Kinyarwanda Automatic Speech Recognition Dataset Dataset Description This dataset contains a sample from the 500 hours of Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition. Dataset Details Language: Kinyarwanda (rw) Task: Automatic Speech Recognition Size: ~500 hours of transcribed speech Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-sample.audioautomatic-speech-recognition1K<n<10K0 likes30 downloads1y agoHugging Face06yigagilbert /kinyarwanda-speech-trimmedgated Kinyarwanda Speech Dataset (Trimmed) This dataset contains processed Kinyarwanda speech data with trimmed audio segments. Dataset Structure The dataset contains two splits: dev_test: 9,263 samples test: 9,265 samples Features Each sample contains: id: Unique identifier for the sample audio: Audio data audio_language: Language of the audio (Kinyarwanda) text: Transcription of the audio prompt: Associated prompt or context duration: Duration of the audio… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/kinyarwanda-speech-trimmed.audioautomatic-speech-recognition10K<n<100K0 likes17 downloads1y agoHugging Face07vysakh25 /kinyarwanda-pastor-snac Kinyarwanda Pastor Uwambaje — SNAC Dataset Single-speaker Kinyarwanda TTS dataset from Pastor Uwambaje YouTube sermons. Audio cleaned with htdemucs (music removal) + resemble-enhance (denoising). Encoded with SNAC 24kHz, 7-token interleaved format. Total uploaded: 13,380 clips at STOI >= 0.80 threshold (~18.4h). STOI Quality Distribution STOI Threshold Clips Hours >= 0.80 (this dataset) 13,225 18.4h >= 0.85 13,058 18.2h >= 0.90 12,488 17.5h >= 0.95 9,734… See the full description on the dataset page: https://huggingface.co/datasets/vysakh25/kinyarwanda-pastor-snac.audiotext-to-speech10K<n<100K0 likes7 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.