datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NigBench-MAMAI-Speech-QA
Voices for Smart Care
A Rural-First Multilingual Voice Dataset for Maternal Health in Nigeria
Voices for Smart Care is a multilingual speech dataset containing real-world maternal and reproductive health questions collected from women across Nigeria. The dataset was created to support the development and evaluation of Automatic Speech Recognition (ASR) and Large Language Models (LLMs) for low-resource African languages in healthcare settings.
Unlike generic speech… See the full description on the dataset page: https://huggingface.co/datasets/intronhealth/NigBench-MAMAI-Speech-QA.synth-qa-taste-codec-chat
Synthetic QA Taste-S Codec Chat
18571 single-turn Traditional Chinese QA utterances with synthesized speech, 21.6 hours of
audio before codec extraction.
Assistant speech is represented as:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
Each text token is followed by its 16 Taste-S FSQ codes (codebooks a..p).
Configurations
default — messages (user question + assistant <SAY> speech), audio, and answer text.
Statistics
Utterances: 18571… See the full description on the dataset page: https://huggingface.co/datasets/yilele/synth-qa-taste-codec-chat.corp23_QA
🗣️ Common Voice 23 — Georgian ASR Quality Control Dataset
Overview
This dataset was created as part of a quality control (QC) process for the Mozilla Common Voice Corpus 23 (Georgian subset).The main goal is to identify corrupted, low-quality, or mislabeled recordings that might have passed through Common Voice validation but are unsuitable for training or evaluation.
🧩 Methodology
User SelectionUnique users were selected from the Common Voice 23 Georgian… See the full description on the dataset page: https://huggingface.co/datasets/psyfreak/corp23_QA.
