CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ssz1111 /SpokenWOZ-Test-Audioaudio1K<n<10K1 likes1k downloads9mo agoHugging Face02Jack-ppkdczgx /SEA-Spoofgated SEA-Spoof Access And License Access requires author approval. Please email the authors before requesting or using the dataset: wu_jinyang@a-star.edu.sg imcc.sg@gmail.com This dataset is released for non-commercial academic research only. Use is restricted to academic institutions and approved research users. Commercial use is not permitted, and this dataset may not be used by commercial companies or for commercial products, services, model training, evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Jack-ppkdczgx/SEA-Spoof.audioaudio-classification100K<n<1M6 likes989 downloads2mo agoHugging Face03mteb /spoken-squad-t2aaudiotext-retrievaln<1K0 likes855 downloads8mo agoHugging Face04Atotti /spoken-multiturn-sft Spoken Multi-turn SFT Japanese Japanese spoken multi-turn SFT dataset generated from kanhatakeyama/AutoMultiTurnByCalm3-22B using CosyVoice2 TTS. Dataset Description This dataset contains Japanese multi-turn SFT (Supervised Fine-Tuning) data with spoken questions. q1: First question (text + audio) a1: First answer (text only) q2: Follow-up question (text + audio) a2: Second answer (text only) Samples ID Q1 Q1 Audio A1 Q2 Q2 Audio A2 0 鉄は強磁性体ですか?… See the full description on the dataset page: https://huggingface.co/datasets/Atotti/spoken-multiturn-sft.audio10K<n<100K0 likes608 downloads9mo agoHugging Face05GSQA /spoken-alpaca-gpt4audio10K<n<100K8 likes514 downloads3y agoHugging Face06abdulahh35 /ANC-Spoof ANC-Spoof: Audio Neural Codec Spoof Dataset Overview ANC-Spoof is a large-scale dataset for studying the robustness of audio deepfake detection (ADD) systems against distortions introduced by neural audio codecs. It pairs original (uncompressed) audio from three established ADD benchmarks with codec-resynthesized versions of the same utterances, produced by eight different neural codecs (including the uncompressed Original version). The dataset is built from:… See the full description on the dataset page: https://huggingface.co/datasets/abdulahh35/ANC-Spoof.audioaudio-classification1M<n<10M0 likes481 downloads2mo agoHugging Face07xxuan-speech /SpoofCelebaudio100K<n<1M0 likes433 downloads2mo agoHugging Face08dianavdavidson /indic-voices-hinglish-nospeakeroverlap-spon3.3-acronyms-fixed2audio100K<n<1M0 likes425 downloads2mo agoHugging Face09QCRI /SpokenNativQA SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs The SpokenNativQA dataset consists of question-answer (QA) pairs, where queries are sourced from real users and answers are manually reviewed and edited. The dataset covers a diverse range of 18 topics that reflect culturally and regionally specific knowledge, as well as everyday queries. These topics include animals, business, clothing, education, events, food and drinks, general knowledge, geography, immigration… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/SpokenNativQA.audioquestion-answering10K<n<100K3 likes412 downloads1y agoHugging Face10MagicHub /multi-stream-spontaneous-conversation-training-datasets_chinese Multi-stream Spontaneous Conversation Training Datasets_Chinese Every data point counts. Dataset Basic Info Dataset Type: ASR Corpus Language: Chinese Audio Parameters: 16 kHz, 16 bits File Format: WAV (PCM) Recording Equipment: Mobile device Dataset Description The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/multi-stream-spontaneous-conversation-training-datasets_chinese.audio1K<n<10K2 likes407 downloads3mo agoHugging Face11AudioLLMs /spoken_squad_testThis dataset is licensed under the terms of the CC-BY-SA-4.0 license. https://github.com/Chia-Hsuan-Lee/Spoken-SQuAD/blob/master/LICENSE.md Author: @michaellee886 @article{li2018spoken, title={Spoken SQuAD: A study of mitigating the impact of speech recognition errors on listening comprehension}, author={Li, Chia-Hsuan and Wu, Szu-Lin and Liu, Chi-Liang and Lee, Hung-yi}, journal={arXiv preprint arXiv:1804.00320}, year={2018} } @article{wang2024audiobench, title={AudioBench: A… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/spoken_squad_test.audio1K<n<10K1 likes394 downloads1y agoHugging Face12slprl /SpokenSwag SpokenSwag We present here SpokenSwag as described in the paper "Slamming: Training a Speech Language Model on One GPU in a Day". This dataset is based on allenai/swag and synthetised with 4 speakers from hexgrad/Kokoro-82M. We show that perfoming DPO over the dataset can really improve performance of Speech Language Models. We encourage you to also see the following resources, for further information: Project Page: https://pages.cs.huji.ac.il/adiyoss-lab/slamming/ Paper:… See the full description on the dataset page: https://huggingface.co/datasets/slprl/SpokenSwag.audioaudio-to-audio10K<n<100K7 likes354 downloads2y agoHugging Face13Loie /SpotSound-Bench SpotSound-Bench: A 'Needle-in-a-Haystack' Evaluation for Audio Temporal Grounding Benchmark Summary SpotSound-Bench is a challenging temporal grounding benchmark designed to evaluate Large Audio-Language Models (ALMs). Existing benchmarks for audio temporal grounding often feature high ratios of target-window duration to full audio clip duration, which fail to simulate real-world scenarios where short events are obscured by dense background sounds. To bridge… See the full description on the dataset page: https://huggingface.co/datasets/Loie/SpotSound-Bench.audion<1K1 likes350 downloads2mo agoHugging Face14dianavdavidson /indic-voices-hinglish-nospeakeroverlap-sponaudio100K<n<1M0 likes312 downloads5mo agoHugging Face15michaelcacioli /Neapolitan-Spoken-Corpus Neapolitan Spoken Corpus (NSC) A corpus of read Neapolitan speech for ASR evaluation, with a validated Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters, metric implementations, per-clip results, and error annotations. This release supersedes the earlier 141-clip single-speaker version of this repository. The earlier release corresponds to Speaker S1 of the present corpus; the old audioData/ and transcripts.csv are replaced by data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/michaelcacioli/Neapolitan-Spoken-Corpus.audioautomatic-speech-recognitionn<1K4 likes308 downloads3mo agoHugging Face16dianavdavidson /indic-voices-hinglish-nospeakeroverlap-spon3audio100K<n<1M0 likes298 downloads3mo agoHugging Face17laudite-ufg /podcasts_spotifyaudio100K<n<1M0 likes292 downloads1y agoHugging Face18dianavdavidson /indic-voices-hinglish-nospeakeroverlap-spon3.1audio100K<n<1M0 likes278 downloads3mo agoHugging Face19MagicDataTech /multi-stream-spontaneous-conversation-training-datasets_chinese Multi-stream Spontaneous Conversation Training Datasets_Chinese Every data point counts. Dataset Basic Info Dataset Type: ASR Corpus Language: Chinese Audio Parameters: 16 kHz, 16 bits File Format: WAV (PCM) Recording Equipment: Mobile device Dataset Description The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicDataTech/multi-stream-spontaneous-conversation-training-datasets_chinese.audio1K<n<10K2 likes275 downloads4mo agoHugging Face20ssz1111 /SpokenWOZ-Train-Audioaudio1K<n<10K0 likes265 downloads9mo agoHugging Face21Atotti /spoken-magpie-ja Spoken-magpie LLMの日本語Instruction Tuning用データllm-jp/magpie-sft-v1.0をCosyVoice2 TTSを使用して音声化した商用利用可能な日本語の音声言語モデルのSFT用データセットです。 ある程度の話者多様性を持つように生成されています。 Respone Audioは500文字以下の場合にのみ生成されています。 NVIDIA H200を10枚を使用しvllmで推論しました。 Samples 最初の50サンプルを掲載します。 ID Instruction Instruction Audio Response Response Audio 0 カボチャを使ったスイーツのレシピをいくつか教えてください。 もちろんです、カボチャを使ったスイーツは秋にぴったりですね。以下にいくつかのレシピをご紹介します。1. カボチャのスフレパウンドケーキ- 材料:カボチャ 200g、生クリーム 50ml、牛乳 50ml、卵 3個、砂糖 100g、薄力粉 70g、バニラエッセンス 少々-… See the full description on the dataset page: https://huggingface.co/datasets/Atotti/spoken-magpie-ja.audiotext-generation100K<n<1M1 likes247 downloads9mo agoHugging Face22DynamicSuperbPrivate /SpokenTermDetection_Tedlium2Train Dataset Card for "SpokenTermDetection_Tedlium2Train" More Information needed audio10K<n<100K0 likes239 downloads3y agoHugging Face23ICTNLP /SpokenVisITSpokenVisIT SpokenVisIT is a real-world visual-speech interaction benchmark built upon VisIT-Bench, designed to evaluate the visual-grounded speech interaction capabilities of omni large multimodal models (LMMs). Our deepest acknowledgment goes to VisIT-Bench — A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use — which collects a diverse set of real-world visual instructions. SpokenVisIT builds on this foundation by converting the textual instructions into spoken… See the full description on the dataset page: https://huggingface.co/datasets/ICTNLP/SpokenVisIT.audion<1K1 likes227 downloads1y agoHugging Face24vendrkat /spokenwoz_dst Dataset: spokenwoz_whisper_dst Short description Prepared SpokenWoZ training dataset adapted for Whisper-style DST (dialog state tracking) tasks. Each example contains an utterance, normalized text, and aligned audio (16 kHz). Key metadata Number of examples: 73,950 Total size on disk: ~9.08 GB Splits: train, validation validation examples: 7,284 Features: text (string): original utterance text normalized_text (string): normalized form of the… See the full description on the dataset page: https://huggingface.co/datasets/vendrkat/spokenwoz_dst.audio100K<n<1M0 likes220 downloads3mo agoHugging Face25dianavdavidson /indic-voices-hinglish-nospeakeroverlap-spon3.2audio100K<n<1M0 likes207 downloads3mo agoHugging Face26sukhdeveyash /partial-spoof-cross-domain-audit-data Partial-Spoof Cross-Domain Audit: detector score outputs Per-utterance and per-frame score outputs of three detectors (MRM, BAM, CFPRF) on PartialSpoof, LlamaPartialSpoof, PartialEdit, and HQ-MPSD. Together with the analysis code they reproduce the reported results of the cross-domain operational audit. An earlier version of this study was submitted to IJCB 2026; that submission was withdrawn and was never published. These arrays support one manuscript, currently in preparation… See the full description on the dataset page: https://huggingface.co/datasets/sukhdeveyash/partial-spoof-cross-domain-audit-data.audio0 likes199 downloads1mo agoHugging Face27blitt /SPoRCgated SPoRC: the Structured Podcast Open Research Corpus (V 1.1) SPoRC is a large multimodal dataset for studying the podcast ecosystem. It contains metadata, full transcripts, speaker-turn-level diarization, speaker-role labels, and acoustic features for over 1.1 million podcast episodes across 228,000 podcasts. Paper: Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus (ACL 2025) Upgrading from version 1.0? Read What changed since version 1.0 first. The file… See the full description on the dataset page: https://huggingface.co/datasets/blitt/SPoRC.tabulartext-classification100M<n<1B26 likes185 downloads2mo agoHugging Face28dianavdavidson /indic-voices-hinglish-nospeakeroverlap-spon3.3audio100K<n<1M0 likes184 downloads3mo agoHugging Face29liepa-project /LIEPA-3_read-spon_16kHzgated LIEPA-3 Bendrinis Garsynas Didysis lietuvių kalbos garsynas (LIEPA-3) – atviras kalbos duomenų rinkinys, skirtas šnekos atpažinimo tikslams ir moksliniams tyrimams. Aprašymas Šis duomenų rinkinys apima LIEPA-3 garsyno read ir spon dalis (fonetinė phon ir dialektų dial dalys neįtrauktos). Duomenys paruošti Whisper mokymui. Pakeitimų istorija 1 versija — 2026-08-25 (20260825) Pradinis pilno garsyno įkėlimas: 6,729,976 įrašai (be P) Kalbėtojų… See the full description on the dataset page: https://huggingface.co/datasets/liepa-project/LIEPA-3_read-spon_16kHz.audioautomatic-speech-recognition1M<n<10M0 likes184 downloads26d agoHugging Face30MagicHub /multi-stream-spontaneous-conversation-training-datasets_english Multi-stream Spontaneous Conversation Training Datasets_English Every data point counts. Dataset Basic Info Dataset Type: ASR Corpus Language: English Audio Parameters: 16 kHz, 16 bits File Format: WAV (PCM) Recording Equipment: Mobile device Dataset Description The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/multi-stream-spontaneous-conversation-training-datasets_english.audio1K<n<10K1 likes179 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.