datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TurkmenSpeech
Turkmen Speech Dataset (ASR)
This dataset contains 251 hours of Turkmen speech audio with transcriptions, intended for training and evaluating Automatic Speech Recognition (ASR) models.
It is one of the largest publicly available Turkmen speech datasets.
Dataset Overview
Property
Value
Total clips
119,847
Total duration
251.86 hours
Sampling rate
16,000 Hz
Language
Turkmen (tk)
Split
train
Each item includes:
audio: waveform + sampling… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenSpeech.NigBench-MAMAI-Speech-QA
Voices for Smart Care
A Rural-First Multilingual Voice Dataset for Maternal Health in Nigeria
Voices for Smart Care is a multilingual speech dataset containing real-world maternal and reproductive health questions collected from women across Nigeria. The dataset was created to support the development and evaluation of Automatic Speech Recognition (ASR) and Large Language Models (LLMs) for low-resource African languages in healthcare settings.
Unlike generic speech… See the full description on the dataset page: https://huggingface.co/datasets/intronhealth/NigBench-MAMAI-Speech-QA.
