CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01QCRI /SpokenNativQA SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs The SpokenNativQA dataset consists of question-answer (QA) pairs, where queries are sourced from real users and answers are manually reviewed and edited. The dataset covers a diverse range of 18 topics that reflect culturally and regionally specific knowledge, as well as everyday queries. These topics include animals, business, clothing, education, events, food and drinks, general knowledge, geography, immigration… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/SpokenNativQA.audioquestion-answering10K<n<100K3 likes412 downloads1y agoHugging Face02ssz1111 /SpokenWOZ-Train-Text What is SpokenWOZ? SpokenWOZ is a large-scale multi-domain speech-text dataset for spoken task-oriented dialogue modeling, which consists of 203k turns, 5.7k dialogues and 249 hours audios from realistic human-to-human spoken conversations. Why SpokenWOZ? The majority of existing TOD datasets are constructed via writing or paraphrasing from annotators rather than being collected from realistic spoken conversations. The written TDO datasets may not be representative of the… See the full description on the dataset page: https://huggingface.co/datasets/ssz1111/SpokenWOZ-Train-Text.text1K<n<10K0 likes399 downloads9mo agoHugging Face03orcarouter /spoken-multihop-rag Spoken Multi-hop QA: ASR Transcripts Across Four English Accents ASR transcriptions of 3,000 multi-hop QA questions, each spoken in four English accents and transcribed with Whisper-large-v3. Released as the data companion to Better Retrieval, Worse Robustness: How Multi-hop RAG Amplifies Upstream ASR Errors (EMNLP 2026, Main Conference). The dataset exists to make one thing cheap to study: what happens to a retrieval pipeline when its query arrives through ASR rather than as… See the full description on the dataset page: https://huggingface.co/datasets/orcarouter/spoken-multihop-rag.textquestion-answering10K<n<100K4 likes338 downloads1mo agoHugging Face04alinet /spoken_squad Dataset Card for Spoken-SQuAD Citation @article{lee2018spoken, title={Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension}, author={Lee, Chia-Hsuan and Wu, Szu-Lin and Liu, Chi-Liang and Lee, Hung-yi}, journal={Proc. Interspeech 2018}, pages={3459--3463}, year={2018} } textquestion-answering10K<n<100K1 likes116 downloads3y agoHugging Face05freococo /myanmar_yes_affirmation_spoken_dataset Myanmar Yes Affirmation Spoken Dataset Creator: freococoLicense: CC0 1.0 (Public Domain)Recommended for: Hugging Face, LLM fine-tuning, NLP researchTested with: Gemini Pro 3.0, ChatGPT 5.0 📖 Dataset Description This dataset contains spoken Burmese expressions that all convey the meaning of "yes" / affirmation. It includes variations across: Formality levels (casual, polite, formal) Speaker gender (male, female, unisex) Contextual usage (friends, shopkeepers… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_yes_affirmation_spoken_dataset.texttext-classificationn<1K0 likes8 downloads9mo agoHugging Face06z050209 /va-spoken-qa-agentvoice-stage2 Stage-2 spoken-QA training corpus (gemma4_talker) 97,697 train / 935 val spoken QA pairs. Questions: REAL user audio from VoiceAssistant-400K (audio_q.*.tar, filenames match input_audio basenames in the manifests). Answers: synthesized in ONE fixed agent voice (LibriSpeech train-clean-100 narrator ref via ResembleAI/Chatterbox; audio_a.*.tar matching assistant_audio). Manifests carry transcripts, answer text, and GLM-4-Voice speech tokens for both sides (glm_in_tokens question /… See the full description on the dataset page: https://huggingface.co/datasets/z050209/va-spoken-qa-agentvoice-stage2.tabularaudio-to-audio10K<n<100K0 likes8 downloads1mo agoHugging Face07angelacao /gloss_to_spoken5textn<1K0 likes6 downloads3y agoHugging Face08atomkwk /srt-SpokenCantoneseToWrittenChinese#Introduction to this dataset This data set is for training llm to translate spoken cantonese srt to written chinese srt(Words in this set is written as simplified chinese characters). Each input and output contain a group of 10 sentances, with a next line character \n between each sentance. textn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.