luganda
Datasets
All datasets matching “luganda”LugandaSoloSpeech_1K
LugandaSoloSpeech1K
1,000+ Hours of single-speaker(s) Unlabeled Luganda Speech Dataset. Perfect for Speech-To-Text / ASR.
Audio quality varies from good to noisy & background music.
Dataset Details
Format: MP3, Mono, 64kbps, 16KHz
Size: 42GB
Data Sources
Radio shows, Youtube
luganda
Luganda Sentences Corpus
A cleaned corpus of Luganda sentences extracted from various publicly available text dumps and corpora.
Motivation
Publicly available text dumps can contain text from languages other than the language they are intended to represent. This can introduce unwanted language contamination into downstream NLP models, even when the original dumps have already undergone cleaning.
This corpus focuses on reducing that contamination by filtering the… See the full description on the dataset page: https://huggingface.co/datasets/sagepond/luganda.Processed-Luganda-SpeechT5-with-SALT-translation-11-7-23
Dataset Card for "Processed-Luganda-SpeechT5-with-SALT-translation-11-7-23"
More Information needed
luganda_callhome_diarization_dataset_MHDPluganda-english-cleaned-v1-splitenglish_luganda
English–Luganda Dataset
This dataset is a reconstruction and cleaning effort by SAGE POND to address noise and quality issues identified in the English–Luganda data derived from the NLLB (No Language Left Behind) corpus.
Purpose
The original NLLB-derived data contains noisy, inconsistent, and incorrectly aligned sentence pairs. This dataset aims to reconstruct and improve the English–Luganda pairs for use in:
Machine translation
Multilingual language-model… See the full description on the dataset page: https://huggingface.co/datasets/sagepond/english_luganda.
