datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kinya-ag-tts
Kinyarwanda Agricultural Text-to-Speech Dataset
In Rwanda, many farmers struggle to access timely, personalized agricultural information. Traditional channels - like radio, TV, and online sources - offer limited reach and interactivity, while extension services and a national call center, staffed by only two agents for over two million farmers, face capacity constraints. To address these gaps, we developed a 24/7 AI-enabled Interactive Voice Response (IVR) tool. Accessible via a… See the full description on the dataset page: https://huggingface.co/datasets/C4IR-RW/kinya-ag-tts.kin_cleaned_commonvoice_rwanda_200hoursRWAVS
AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis
Susan Liang, Chao Huang, Yapeng Tian, Anurag Kumar, Chenliang Xu
RWAVS Dataset
We provide the Real-World Audio-Visual Scene (RWAVS) Dataset.
The dataset can be downloaded from this Hugging Face repository.
After you download the dataset, you can decompress the RWAVS_Release.zip.
unzip RWAVS_Release.zip
cd release/
The data is organized with the following directory structure.
./release/
├── 1
│… See the full description on the dataset page: https://huggingface.co/datasets/susanliang/RWAVS.rwandan_kinyarwanda_nonstandard_speech_v1.0This dataset provides 61.7 hours of Kinyarwanda speech recordings (14,739 samples) from 61 Rwandan speakers living with speech impairments. The participants represent a limited diversity of speech patterns, mostly stuttering, and a few examples of Dysarthria, Dysphonia, and Phonological disorders.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or phrase level. All speech recordings of this datasets have been… See the full description on the dataset page: https://huggingface.co/datasets/cdli/rwandan_kinyarwanda_nonstandard_speech_v1.0.KinyaWhisperDataset
Kinyarwanda Spoken Words Dataset
This dataset contains 102 short audio samples of spoken Kinyarwanda words, each labeled with its corresponding transcription. It is designed for training, evaluating, and experimenting with Automatic Speech Recognition (ASR) models in low-resource settings.
Structure
audio/: Contains 102 .wav files (mono, 16kHz)
transcripts.txt: Tab-separated transcription file (e.g., 001.wav\tmuraho)
manifest.jsonl: JSONL file with audio paths and text… See the full description on the dataset page: https://huggingface.co/datasets/benax-rw/KinyaWhisperDataset.mlcs_rw_pitchrwandan_english_nonstandard_speech_v1.0This dataset provides 32.7 hours of English speech recordings (7,592 samples) from 44 Rwandan speakers living with speech impairments. The participants represent a limited diversity of speech patterns, mostly stuttering, and a few examples of Dysarthria, Dysphonia, and Phonological disorders.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or phrase level. All speech recordings of this datasets have been… See the full description on the dataset page: https://huggingface.co/datasets/cdli/rwandan_english_nonstandard_speech_v1.0.rw-tts-dataset
Rw Tts Dataset
Dataset Description
Kinyarwanda (rw) text-to-speech dataset. Studio-recorded read speech aligned with transcriptions, collected in Rwanda.
Languages
Language: Kinyarwanda (rw)
BCP-47: rw
Source tag
rw — identifies the origin of each sample in the source column.
Dataset Structure
Column
Type
Description
audio
Audio
Raw WAV audio at original recording frequency
text
string
Transcription of the spoken content… See the full description on the dataset page: https://huggingface.co/datasets/Bateesa/rw-tts-dataset.rwc-poprwords
Dataset Card for rwords
Аудиофайлы выговаривания звука "Р"
Dataset Details
Dataset Description
В датасет собраны аудиофайлы отдельных слов, в которых "Р" выговаривается хорошо или плохо.
Файлы организованы в два каталога по классам: good/ (1162 файла) и notgood/ (1162 файла).
Curated by: Павел Рудич
Funded by: Фонд содействия инновациям
Language(s) (NLP): Russian
License: MIT
Uses
Может быть использован для изучения особенностей произношения звука… See the full description on the dataset page: https://huggingface.co/datasets/dysata/rwords.eval_rw_2bengali-tts-combined
