datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twi-health-asr-gemini-500hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Health Speech Dataset Gemini (500 hours)
A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken
languages, sourced from publicly available video content on health and wellness.
Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-health-asr-gemini-500hrs.twi-health-asr
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Health Speech Dataset
A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken languages,
sourced from publicly available video content on health and wellness.
Created by Mich-Seth Owusu and… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-health-asr.twi-health-asr-gemini-500hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Health Speech Dataset Gemini (500 hours)
A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken
languages, sourced from publicly available video content on health and wellness.
Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-health-asr-gemini-500hrs.twi-health-asr-gemini-500hrs-ipa
Twi Health Speech — Audio, Transcript and IPA
Twi health-domain speech with both a written transcript and an IPA phoneme sequence read off the audio by ASR. Built from ghananlpcommunity/twi-health-asr-gemini-500hrs by adding the IPA column.
from datasets import load_dataset
ds = load_dataset("ghananlpcommunity/twi-health-asr-gemini-500hrs-ipa", split="train")
ds[0]["audio"] # decoded waveform, 16 kHz
ds[0]["transcription"] # transcript
ds[0]["ipa"]… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-health-asr-gemini-500hrs-ipa.voice-of-care-health-dataset
Voice of Care AI for Global Health Benchmark Dataset
Overview
This dataset contains spoken Hausa Health datasets with rich annotations covering emotion, intent, speaker demographics, and dialect variation, intended for speech and NLP research.
Dataset Summary
Property
Details
Language
Hausa
Modality
Audio + Text
Task(s)
e.g. Speech Recognition, Emotion Detection, Dialect Identification
Version
1.0.0
🛠️ Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Data-Science-Nigeria/voice-of-care-health-dataset.weha-health-benchmark-audio
Weha Health Benchmark Audio
Consented, de-identified voice recordings collected for benchmarking
speech-to-text engines as part of the Sahara CodeSwitch Africa Challenge
(Intron Health), Health category submission "Weha Health."
Contents
60 short simulated health-triage utterances across 4 languages, recorded
by the Weha Health team, each naturally code-switched with English
(except the Yoruba subset, which is monolingual Yoruba):
Yoruba (yo): 20 clips
Nigerian… See the full description on the dataset page: https://huggingface.co/datasets/techwithnel/weha-health-benchmark-audio.ASR_mental_health_luganda_dataset
Luganda ASR Mental Health Dataset
Dataset Description
This dataset contains Luganda speech recordings with corresponding transcriptions focused on mental health conversations. The dataset follows the Common Voice structure and is designed for automatic speech recognition research in low-resource African languages.
Dataset Summary
The Luganda ASR Dataset is a specialized speech recognition corpus for Luganda (ISO 639-1: lg), primarily focused on… See the full description on the dataset page: https://huggingface.co/datasets/africanobyamugisha/ASR_mental_health_luganda_dataset.
