CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01smgjch /meow-10k Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology. Dataset Summary Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/smgjch/meow-10k.audio10K<n<100K3 likes1.5k downloads5mo agoHugging Face02anonymous2222 /Sympatheia-18k Sympatheia-18k Sympatheia-18k is an emotion-aware spoken dialogue dataset for empathetic speech synthesis research. It contains 18,000 query–response pairs across 12 emotion categories, each accompanied by synthesized audio and text transcripts. Dataset Structure Subset Unique Queries Responses Description Emotional 8,400 train / 3,600 eval 8,400 train / 3,600 eval Emotional queries with emotionally-matched responses Neutral 350 train / 150 eval 4,200 train /… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2222/Sympatheia-18k.audioaudio-to-audio10K<n<100K0 likes477 downloads5mo agoHugging Face03ground-truth /multichannel-meetings-10h GroundTruth Multi-Channel Meeting Audio Dataset (10h) Summary This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant. Each meeting includes: One full meeting recording (room microphone) Individual close-talk recordings for each participant (one file per speaker) Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.audioautomatic-speech-recognitionn<1K1 likes287 downloads5mo agoHugging Face04susameddin /Sympatheia-18k Sympatheia-18k Sympatheia-18k is an emotion-aware spoken dialogue dataset for empathetic speech synthesis research. It contains 18,000 query–response pairs across 12 emotion categories, each accompanied by synthesized audio and text transcripts. Dataset Structure Subset Unique Queries Responses Description Emotional 8,400 train / 3,600 eval 8,400 train / 3,600 eval Emotional queries with emotionally-matched responses Neutral 350 train / 150 eval 4,200… See the full description on the dataset page: https://huggingface.co/datasets/susameddin/Sympatheia-18k.audioaudio-to-audio10K<n<100K2 likes206 downloads4mo agoHugging Face05zcs15 /SVHalluc SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models TL;DR: Speech content is not necessarily visual evidence. SVHalluc is the first benchmark that tests whether audio-visual large language models (AV-LLMs) can distinguish what is said from what is actually seen. SVHalluc was introduced in “SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models,” CVPR 2026. Project page Paper (arXiv) Code… See the full description on the dataset page: https://huggingface.co/datasets/zcs15/SVHalluc.textvisual-question-answering1K<n<10K0 likes205 downloads2mo agoHugging Face06aseth125 /audio-hallucination-attack Audio Hallucination Attacks (AHA) Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models" It contains two subsets: AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training Audio Files The audio files are provided as compressed archives in this repository: File Contents Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.audioaudio-classification100K<n<1M2 likes191 downloads6mo agoHugging Face07mcp-tool-shop /jam-actions-v1 jam-actions-v1 Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) · Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE · Source repo: mcp-tool-shop-org/ai-jam-sessions The successor to jam-actions-v0. Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from what the tools return — and it exists in its current shape because, seven training runs in a row, the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.texttext-generationn<1K0 likes188 downloads15d agoHugging Face08mcp-tool-shop /jam-actions-v1-probe jam-actions-v1-probe Schema: jam-actions-v1-probe/1.0.0 · Records: 24, all split: test · Evaluation only · Companion to: jam-actions-v1 Why it exists An adapter trained on an earlier version of the corpus scored 47/54 on held-out acoustic takes. Its completions, which state the comparison before the label, showed that it wrote against a 50-cent gate whenever it saw a minus sign — and negative cents occurred in exactly one class of that corpus. The main split could… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1-probe.textothern<1K0 likes158 downloads14d agoHugging Face09nekoyama12 /Music Made by herza For APP.t audion<1K0 likes156 downloads4mo agoHugging Face10tterumiimurett1 /agentic-asrgated Agentic ASR Public consolidated audio and ASR result dataset for the OSWorld and WildClawBench benchmark families. Layout osworld/: synthetic raw/colloquial speech, human recordings, DNS-noise pairs, task images, ASR results, and reports. wildclawbench/: 60 formal colloquialized prompts, synthetic speech, 20 synthetic ASR condition tables, and ten-participant human recordings. task0_template derivatives are excluded. metadata/conditions.jsonl: model, variant… See the full description on the dataset page: https://huggingface.co/datasets/tterumiimurett1/agentic-asr.audio10K<n<100K0 likes139 downloads4d agoHugging Face11yunqi1766 /voice-code-bench VoiceCodeBench VoiceCodeBench is a test-only benchmark for evaluating whether automatic speech recognition (ASR) systems preserve exact structured values in English workplace speech. Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition The benchmark targets cases where a transcript is software input: callback numbers, email addresses, command-line flags, file paths, URLs, account identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/yunqi1766/voice-code-bench.audioautomatic-speech-recognitionn<1K1 likes123 downloads2mo agoHugging Face12m-a-p /Chords1217gatedaudio1K<n<10K5 likes93 downloads1y agoHugging Face13tsdocode /open-vi-dialog-synthetic-100h OpenDialog Vietnamese Synthetic Dialogue 100h Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments. 12,000 chunks 30 seconds per chunk 100.0 hours total Each item contains S1/S2 speaker labels, turn timings, target text, relationship, pronouns, environment, topic, mood, and source reference IDs. Audio renderer: vLLM-Omni VoxCPM2 Audio format: mono WAV, 48 kHz, 30 seconds per chunk This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.audiotext-to-speech10K<n<100K0 likes92 downloads1mo agoHugging Face14AudioCC-Lab /PICSAFEv1 Speech Quality Test Labels PICSAFEv1 is a multi-source annotated test dataset for evaluating speech quality assessment and audio data filtering methods. It contains 10,728 audio samples drawn from 14 source datasets, with annotations from a vocabulary of 33 tags. These tags describe recording provenance, speech styles, speaking rate and pitch, speaker attributes, noise, reverberation, distortion, and transcript errors. These annotations support benchmarking quality metrics and… See the full description on the dataset page: https://huggingface.co/datasets/AudioCC-Lab/PICSAFEv1.textaudio-classification10K<n<100K0 likes86 downloads8h agoHugging Face15isabeth /rgad-crosslingual-tts-10h RGAD Cross-Lingual TTS 10h This is a 10-hour cross-lingual TTS dataset for prompt-conditioned Chinese TTS fine-tuning. Format The dataset contains: train.jsonl dev.jsonl metadata.csv audio/prompts/*.wav audio/targets/*.wav Each JSONL row has this format: {"id":"sample_000001","prompt_wav":"audio/prompts/sample_000001.wav","target_wav":"audio/targets/sample_000001.wav","text":"中文目标文本。","prompt_language":"en-US","target_language":"zh-CN"… See the full description on the dataset page: https://huggingface.co/datasets/isabeth/rgad-crosslingual-tts-10h.audiotext-to-speech1K<n<10K1 likes79 downloads4mo agoHugging Face16Duckyle /meow-10k Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology. Dataset Summary Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/Duckyle/meow-10k.audio10K<n<100K0 likes75 downloads4mo agoHugging Face17MagicLuke /fdb-v1-outputs-v1gated Full-Duplex-Bench v1.0 model outputs (fdb-v1-outputs-v1) 7997 model responses = 11 benchmark runs × the 727 stimuli of Full-Duplex-Bench v1.0 (pause handling, backchannel, smooth turn-taking, user interruption). The runs cover 7 systems; several differ only in voice prompt, prompting regime or weights, which is the point — those are controlled pairs. For every stimulus and run you get the model's own reply channel as lossless FLAC — time-synchronous with the stimulus, so… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/fdb-v1-outputs-v1.audio1K<n<10K0 likes69 downloads28d agoHugging Face18playwithmino /aishell1mix-ver2-n100-per-mix AISHELL-1 Mix ver2 — 100 clips per mix This is AISHELL-1 Mix ver2, not ver1. 8 kHz mono test subset: 100 mixtures per speaker count (N=1\ldots5) (50 mix_clean + 50 mix_both each) → 500 clips. Derived from the local aishell1mix_ver2 test SCPs (data/scp/scp_aishell1mix_ver2). Includes mixture + oracle speaker stems and transcripts. Split Count 1mix / 2mix / 3mix / 4mix / 5mix 100 each clean / both 250 each Files manifests/test.jsonl —… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/aishell1mix-ver2-n100-per-mix.audioaudio-to-audion<1K0 likes67 downloads24d agoHugging Face19ODYSSEYAILABS /odyssey-v0.5.1-16h-eval-preview-public 🇿🇦 Odyssey V0.5.1 — 16H Evaluation Preview (Code-Switched ASR) Odyssey AI Labs presents a 16-hour Enterprise Evaluation Preview of the V0.5 corpus. This release is designed to surface the hardest South African ASR realities directly: native urban code-switching, multi-speaker turn-taking, and overlapping conversational speech. A structured public preview of the Odyssey SA Voice Corpus, designed for researchers, data buyers, and speech teams evaluating multilingual South African… See the full description on the dataset page: https://huggingface.co/datasets/ODYSSEYAILABS/odyssey-v0.5.1-16h-eval-preview-public.tabularn<1K0 likes59 downloads5mo agoHugging Face20Weisiqing123 /ONOTE ONOTE: Omnimodal Notation Objective Topology Examination ONOTE is a large-scale, omnimodal benchmark designed to evaluate music understanding across three major notation systems: Western Staff, Jianpu (Numbered Notation), and Guitar Tablature. 📂 Dataset Structure The dataset is organized into two primary sub-directories based on the notation and instrument type: 1. pitch_Jianpu_dataset (Staff & Numbered Notation) This sub-folder focuses on Western staff and… See the full description on the dataset page: https://huggingface.co/datasets/Weisiqing123/ONOTE.audioimage-to-text1K<n<10K1 likes57 downloads6mo agoHugging Face21MagicLuke /ifbench-conversations-v1gated IF-Bench conversations (ifbench-conversations-v1) 3000 synthetic full-duplex spoken conversations: 15 examiner configurations (9 speech models), each holding the same 200-task set of Full-Duplex-Bench v2 staged scenarios as the examiner (the model under study — it carries a role, a topic and four goals to hit in order) against a PersonaPlex-7B examinee that is never told the topic. Per dialogue you get both channels as lossless mono FLAC, the exact prompts/voices/sampling… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/ifbench-conversations-v1.audio1K<n<10K0 likes57 downloads28d agoHugging Face22Rakancorle1 /hans-10k Hans-10K · DPO recipe for the audio-visual Clever Hans DPO training data accompanying the paper When Vision Speaks for Sound. Like the original Clever Hans 🐎 — the horse that looked like he could do arithmetic but was actually reading his trainer's body language — video-capable MLLMs often look like they can hear: they answer audio questions by reading visual cues and never verifying the audio stream. Hans-10K is the 10,383-sample best-recipe preference-pair dataset that cures this… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-10k.audioaudio-classification10K<n<100K0 likes52 downloads4mo agoHugging Face23Rakancorle1 /thud-eval THUD-Eval · audio-visual Clever Hans benchmark Evaluation benchmark accompanying the paper When Vision Speaks for Sound. This dataset probes the audio-visual Clever Hans effect — the tendency of video-capable MLLMs to appear to listen while really just reading visual cues. We test the same source clips under three audio interventions: Task Intervention What it tests sync audio temporally shifted (early / delay) Can the model detect a time offset? mute audio replaced… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/thud-eval.audion<1K0 likes41 downloads4mo agoHugging Face24ChristophSchuhmann /advanced-soundscapes-stage-1 Advanced Soundscapes Stage 1 — Raw Components This dataset contains Stage 1 output from the LAION Universal Audio Annotation Pipeline (UAAP) data generation plan. Contents 0 shard(s) containing 0 soundscape recipes with raw audio components Each soundscape row includes: recipe.json — full recipe with timeline, events, loudness, speaker IDs, overlap/density settings spkN.flac / spkN.json — raw speech components + full source metadata musicN.flac / musicN.json —… See the full description on the dataset page: https://huggingface.co/datasets/ChristophSchuhmann/advanced-soundscapes-stage-1.tabularaudio-classificationn<1K0 likes32 downloads3mo agoHugging Face25Rakancorle1 /vggsync-3k VGGSync-3K · out-of-domain audio-visual sync benchmark Out-of-domain evaluation set used in the paper When Vision Speaks for Sound. Derived from VGGSoundSync, this 3,000-clip slice tests whether a video-capable MLLM can detect audio temporal offsets on everyday sound events outside the THUD in-domain training distribution. Each item is one VGGSound clip in one of three conditions: Condition Count gt_synced gt_direction gt_offset_sec Audio aligned (no shift) 1,000 true… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/vggsync-3k.audioaudio-classification1K<n<10K0 likes28 downloads4mo agoHugging Face26eureka1500 /IFAO-lalmtext100K<n<1M0 likes27 downloads5mo agoHugging Face27eureka1500 /CaptionStew10M-Qwen3Omnitext100K<n<1M0 likes27 downloads5mo agoHugging Face28Rakancorle1 /hans-sft-4k Hans-SFT-4K · SFT recipe for the audio-visual Clever Hans Supervised fine-tuning (SFT) data accompanying the paper When Vision Speaks for Sound. Like the original Clever Hans — the horse that looked like he could do arithmetic but was actually reading his trainer's body language — video-capable MLLMs often look like they can hear: they answer audio questions by reading visual cues and never verifying the audio stream. Hans-SFT-4K is the 3,834-sample SFT mix that teaches models to… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-sft-4k.audioaudio-classification1K<n<10K1 likes24 downloads4mo agoHugging Face29ODYSSEYAILABS /odyssey-v0.5.1-16h-eval-audio-gatedgated 🔒 Odyssey V0.5.1 — 16H Gated Audio Evaluation Repo This repository contains the gated audio companion to the public Odyssey evaluation preview. All requests are reviewed manually. Public Preview Repository For transcripts, schema preview, and metadata inspection, see: ODYSSEYAILABS/odyssey-v0.5.1-16h-eval-preview-public About Odyssey AI Labs Odyssey AI Labs builds premium African data infrastructure for next-generation AI systems. audion<1K0 likes20 downloads5mo agoHugging Face30Dramazy /meow-10k Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology. Dataset Summary Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/Dramazy/meow-10k.audio10K<n<100K0 likes19 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.