datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
luganda_bible_audio_100Hrs
Luganda Bible Audio 100Hrs
Audio split by chapters of Old & New Testament & Transcripts
Format: 64kps, MP3
Requires further splitting of the audio into smaller chunks/splits (https://github.com/facebookresearch/fairseq/tree/main/examples/mms/data_prep)
Audio sourced from https://www.faithcomesbyhearing.com/
ASR_mental_health_luganda_dataset
Luganda ASR Mental Health Dataset
Dataset Description
This dataset contains Luganda speech recordings with corresponding transcriptions focused on mental health conversations. The dataset follows the Common Voice structure and is designed for automatic speech recognition research in low-resource African languages.
Dataset Summary
The Luganda ASR Dataset is a specialized speech recognition corpus for Luganda (ISO 639-1: lg), primarily focused on… See the full description on the dataset page: https://huggingface.co/datasets/africanobyamugisha/ASR_mental_health_luganda_dataset.
