datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kenyan_swahili_nonstandard_speech_v1.0This dataset provides 32.5 hours of Swahili speech recordings (5,535 samples) from 52 Kenyan speakers living with speech impairments. The participants represent a diversity of progressive, acquired and congenital aetiologies, including cerebral palsy, Parkinson's disease, multiple sclerosis, autism spectrum disorder, Down syndrome, stroke and stuttering.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or… See the full description on the dataset page: https://huggingface.co/datasets/cdli/kenyan_swahili_nonstandard_speech_v1.0.SwahiliMultimodalSarcasm
Dataset Card for Dataset Name
This dataset is a Swahili multimodal, text and image dataset intended for Sarcasm NLP.
This is a link to the github page: https://github.com/ekariba/SyntheticDataset/blob/main/Dataset_generation.ipynb
Dataset Details
Dataset Description
Curated by: Eugene Kariba
Language(s) (NLP): Swahili
Uses
Multimodal Swahili Sarcasm for environmental themes
Direct Use
To be used for Multimodal Sarcasm analysis in… See the full description on the dataset page: https://huggingface.co/datasets/EKariba/SwahiliMultimodalSarcasm.afrivoice-swahili-agriculture-subset
Dataset Card for the image text and voice dataset
Dataset Description
Subset of Afrivoice dataset:
approx. 100 hours of train (sampled, stratified)
Full dev + test from original repo (DigitalUmuganda/Afrivoice_Swahili)
Audio: .webm
Includes images + transcriptions from original repo (DigitalUmuganda/Afrivoice_Swahili)
Includes metadata csv
License
CC-BY-4.0 (derived from DigitalUmuganda/Afrivoice_Swahili)
Notes
Train split sampled using… See the full description on the dataset page: https://huggingface.co/datasets/egirma/afrivoice-swahili-agriculture-subset.swahili-data
