datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SpokenNativQA
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
The SpokenNativQA dataset consists of question-answer (QA) pairs, where queries are sourced from real users and answers are manually reviewed and edited. The dataset covers a diverse range of 18 topics that reflect culturally and regionally specific knowledge, as well as everyday queries. These topics include animals, business, clothing, education, events, food and drinks, general knowledge, geography, immigration… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/SpokenNativQA.SpokenWOZ-Train-Text
What is SpokenWOZ?
SpokenWOZ is a large-scale multi-domain speech-text dataset for spoken task-oriented dialogue modeling, which consists of 203k turns, 5.7k dialogues and 249 hours audios from realistic human-to-human spoken conversations.
Why SpokenWOZ?
The majority of existing TOD datasets are constructed via writing or paraphrasing from annotators rather than being collected from realistic spoken conversations. The written TDO datasets may not be representative of the… See the full description on the dataset page: https://huggingface.co/datasets/ssz1111/SpokenWOZ-Train-Text.spoken-multihop-rag
Spoken Multi-hop QA: ASR Transcripts Across Four English Accents
ASR transcriptions of 3,000 multi-hop QA questions, each spoken in four
English accents and transcribed with Whisper-large-v3. Released as the
data companion to Better Retrieval, Worse Robustness: How Multi-hop RAG
Amplifies Upstream ASR Errors
(EMNLP 2026, Main Conference).
The dataset exists to make one thing cheap to study: what happens to a
retrieval pipeline when its query arrives through ASR rather than as… See the full description on the dataset page: https://huggingface.co/datasets/orcarouter/spoken-multihop-rag.spoken_squad
Dataset Card for Spoken-SQuAD
Citation
@article{lee2018spoken,
title={Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension},
author={Lee, Chia-Hsuan and Wu, Szu-Lin and Liu, Chi-Liang and Lee, Hung-yi},
journal={Proc. Interspeech 2018},
pages={3459--3463},
year={2018}
}
myanmar_yes_affirmation_spoken_dataset
Myanmar Yes Affirmation Spoken Dataset
Creator: freococoLicense: CC0 1.0 (Public Domain)Recommended for: Hugging Face, LLM fine-tuning, NLP researchTested with: Gemini Pro 3.0, ChatGPT 5.0
📖 Dataset Description
This dataset contains spoken Burmese expressions that all convey the meaning of "yes" / affirmation. It includes variations across:
Formality levels (casual, polite, formal)
Speaker gender (male, female, unisex)
Contextual usage (friends, shopkeepers… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_yes_affirmation_spoken_dataset.va-spoken-qa-agentvoice-stage2
Stage-2 spoken-QA training corpus (gemma4_talker)
97,697 train / 935 val spoken QA pairs. Questions: REAL user audio from
VoiceAssistant-400K (audio_q.*.tar, filenames match input_audio basenames in the
manifests). Answers: synthesized in ONE fixed agent voice (LibriSpeech
train-clean-100 narrator ref via ResembleAI/Chatterbox; audio_a.*.tar matching
assistant_audio). Manifests carry transcripts, answer text, and GLM-4-Voice speech
tokens for both sides (glm_in_tokens question /… See the full description on the dataset page: https://huggingface.co/datasets/z050209/va-spoken-qa-agentvoice-stage2.gloss_to_spoken5srt-SpokenCantoneseToWrittenChinese#Introduction to this dataset
This data set is for training llm to translate spoken cantonese srt to written chinese srt(Words in this set is written as simplified chinese characters).
Each input and output contain a group of 10 sentances, with a next line character \n between each sentance.
