Bateesa/tobydata-tts-dataset
Tobydata Tts Dataset Dataset Description Luganda TTS dataset (Toby-data) collected by TericLab. Contains read speech in Luganda, primarily on tailoring, fashion, and vocational training topics. Recorded via mobile application. Languages Language: Luganda (lg) BCP-47: lg Source tag tobydata — value of the source column in every row. Dataset Structure Column Type Description audio Audio Raw WAV audio at… See the full description on the dataset page: https://huggingface.co/datasets/Bateesa/tobydata-tts-dataset.
Tobydata Tts Dataset
Dataset Description
Luganda TTS dataset (Toby-data) collected by TericLab. Contains read speech in Luganda, primarily on tailoring, fashion, and vocational training topics. Recorded via mobile application.
Languages
- Language: Luganda (lg)
- BCP-47:
lg
Source tag
tobydata — value of the source column in every row.
Dataset Structure
Splits
Related Work
Prepared in the context of East African speech technology research. See also the Sunbird AI SALT dataset — a multi-language parallel corpus covering Luganda, Acholi, Swahili, Runyankole, Lugbara, and Ateso, used as a benchmark for TTS/ASR in East Africa.
License
Citation
@dataset{tobydata_tts,
author = {Bateesa},
title = {Tobydata Tts Dataset},
year = {2025},
publisher = {HuggingFace},
url = {https://huggingface.co/datasets/Bateesa/tobydata-tts-dataset}
}