CoolFace
Datasetpublic

Bateesa/tobydata-tts-dataset

Tobydata Tts Dataset Dataset Description Luganda TTS dataset (Toby-data) collected by TericLab. Contains read speech in Luganda, primarily on tailoring, fashion, and vocational training topics. Recorded via mobile application. Languages Language: Luganda (lg) BCP-47: lg Source tag tobydata — value of the source column in every row. Dataset Structure Column Type Description audio Audio Raw WAV audio at… See the full description on the dataset page: https://huggingface.co/datasets/Bateesa/tobydata-tts-dataset.

sourceHugging Facecc-by-4.0updated 7mo agoView on Hugging Face
1likes21downloads
Dataset Card

Tobydata Tts Dataset

Dataset Description

Luganda TTS dataset (Toby-data) collected by TericLab. Contains read speech in Luganda, primarily on tailoring, fashion, and vocational training topics. Recorded via mobile application.

Languages

  • —Language: Luganda (lg)
  • —BCP-47: lg

Source tag

tobydata — value of the source column in every row.

Dataset Structure

ColumnTypeDescription
audioAudioRaw WAV audio at original recording frequency
textstringTranscription of the spoken content
frequencyintSample rate in Hz of the audio
sourcestringOrigin label (tobydata)

Splits

SplitSamples
train2682

Related Work

Prepared in the context of East African speech technology research. See also the Sunbird AI SALT dataset — a multi-language parallel corpus covering Luganda, Acholi, Swahili, Runyankole, Lugbara, and Ateso, used as a benchmark for TTS/ASR in East Africa.

License

CC BY 4.0

Citation

bibtex
@dataset{tobydata_tts,
  author    = {Bateesa},
  title     = {Tobydata Tts Dataset},
  year      = {2025},
  publisher = {HuggingFace},
  url       = {https://huggingface.co/datasets/Bateesa/tobydata-tts-dataset}
}