luganda
Datasets
All datasets matching “luganda”Processed-Luganda-SpeechT5-with-SALT-translation-11-7-23
Dataset Card for "Processed-Luganda-SpeechT5-with-SALT-translation-11-7-23"
More Information needed
luganda
Luganda Sentences Corpus
A cleaned corpus of Luganda sentences extracted from various publicly available text dumps and corpora.
Motivation
Publicly available text dumps can contain text from languages other than the language they are intended to represent. This can introduce unwanted language contamination into downstream NLP models, even when the original dumps have already undergone cleaning.
This corpus focuses on reducing that contamination by filtering the… See the full description on the dataset page: https://huggingface.co/datasets/sagepond/luganda.LugandaSoloSpeech_1K
LugandaSoloSpeech1K
1,000+ Hours of single-speaker(s) Unlabeled Luganda Speech Dataset. Perfect for Speech-To-Text / ASR.
Audio quality varies from good to noisy & background music.
Dataset Details
Format: MP3, Mono, 64kbps, 16KHz
Size: 42GB
Data Sources
Radio shows, Youtube
luganda_callhome_diarization_dataset_MHDPluganda-english-cleaned-v1-splitSynthetic_Luganda_VITS_22.5k
Dataset Card for "Synthetic_Luganda_VITS_22.5k"
More Information needed
