CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RVtech /Audio2Tool Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1 1 Rivian & Volkswagen Technologies &nbsp;·&nbsp; ∗ equal contribution &nbsp;·&nbsp; ∗∗ corresponding author &nbsp;·&nbsp; † equal contribution 📄 Project page / demo: https://audio2tool.github.io/ 📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool ✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.audioautomatic-speech-recognition10K<n<100K2 likes6k downloads3mo agoHugging Face02cm2435-new /gdpval_preference_rubricsaudion<1K0 likes1.6k downloads5mo agoHugging Face03pollen-robotics /microduck-emotions Microduck Emotions A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.audioroboticsn<1K6 likes952 downloads18d agoHugging Face04RyanWW /XModBenchXModBench Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models 🎉 Accepted at ICLR 2026 What is XModBench? XModBench is the first tri-modal (audio / vision / text) multiple-choice QA benchmark explicitly designed to measure cross-modal consistency — does an omni-language model give the same correct answer when the same semantic content is presented in different modalities? Each item is a 4-choice question with a <context>… See the full description on the dataset page: https://huggingface.co/datasets/RyanWW/XModBench.audiomultiple-choice10K<n<100K3 likes563 downloads4mo agoHugging Face05rookie9 /MMAG MMAG: A Multi‑Control Mixed Audio Generation Benchmark MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment. Dataset Structure The dataset is organized into three subsets, each targeting a specific… See the full description on the dataset page: https://huggingface.co/datasets/rookie9/MMAG.audiotext-to-audio1K<n<10K0 likes365 downloads2mo agoHugging Face06yuanzhuyun /asr-reference-set-eval-temp Temporary ASR evaluation audio Temporary public audio files used for hosted ASR evaluation. audio1K<n<10K0 likes357 downloads2mo agoHugging Face07AnnoymousNeurLPSsubmit /RAIL RAIL Audio Benchmark This folder is generated for direct Hugging Face dataset upload. Each row uses relative audio paths rooted at this repository folder. NeurIPS / Croissant Metadata metadata.json is a Croissant-style metadata file with core fields and minimal RAI fields. metadata.json includes the Hugging Face dataset URL, CC-BY-4.0 license URL, checksums, and RAI fields. build_summary.json contains build counts and skipped-source diagnostics. Schema id:… See the full description on the dataset page: https://huggingface.co/datasets/AnnoymousNeurLPSsubmit/RAIL.audio10K<n<100K1 likes305 downloads5mo agoHugging Face08rajjanardhan00 /Seamless_Dummy_Dataset_Fixed_3 MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. audioquestion-answeringn<1K0 likes294 downloads1y agoHugging Face09RUIH /SlideASR-Benchaudio1K<n<10K0 likes207 downloads1y agoHugging Face10pre-view /CS50-rawaudio10K<n<100K0 likes161 downloads2y agoHugging Face11rorosese /my-voxtral-datasetaudion<1K0 likes152 downloads1y agoHugging Face12YijiaFan /Resource2Skill Resource2Skill — Skill Library Executable skill libraries for the Resource2Skill runtime: reusable, structured skills that a software agent browses, inspects, and composes to operate real tools (Web, PowerPoint, Excel, Blender, and REAPER-style audio) and produce artifacts. This dataset is the skill data half of the project; the runnable runtime, MCP servers, and CLI live in the code repository. Contents skills_wiki/ Structured wiki entries used for runtime… See the full description on the dataset page: https://huggingface.co/datasets/YijiaFan/Resource2Skill.imagen<1K0 likes135 downloads3mo agoHugging Face13sixf0ur /ableton-racks-labeled Dataset Card for Ableton Live XML Presets Dataset Summary This dataset consists of Ableton Live XML presets paired with synthetic, multi-perspective natural language descriptions and Chain-of-Thought (CoT) reasoning. It is designed for fine-tuning Large Language Models (LLMs) to perform text-to-preset generation, sound design analysis, and audio parameter reasoning. Each entry provides three distinct prompt styles (layman, technical, and emotional) alongside a… See the full description on the dataset page: https://huggingface.co/datasets/sixf0ur/ableton-racks-labeled.textn<1K1 likes130 downloads2mo agoHugging Face14rodenhhh /ContextTTS_dataset ContextTTS Evaluation Dataset This is the official evaluation dataset for the paper "[ContextTTS Eval: A Benchmark for Evaluating Long-Form Contextual Expressive Text-to-Speech]". It is designed to evaluate the performance of multi-modal speech synthesis, specifically focusing on context-aware prosody and timbre consistency in Chinese conversations and audiobooks. Dataset Summary The dataset consists of high-quality Chinese audio-text pairs, organized into three distinct… See the full description on the dataset page: https://huggingface.co/datasets/rodenhhh/ContextTTS_dataset.audiotext-to-speechn<1K0 likes120 downloads6mo agoHugging Face15r-labs /kencorpus_sw_culture KenCorpus Swahili Culture Subset A filtered subset of Kencorpus/KenCorpus_audio, containing only rows where language=Swahili and genre=Culture (37 clips). Audio files are in audio/, indexed by kencorpus_sw_culture.jsonl with path and duration fields, following the layout of kyutai/DailyTalkContiguous. audion<1K0 likes104 downloads1mo agoHugging Face16vocalcoachbench /vocalcoachbench-review VocalCoachBench VocalCoachBench is a singing-audio benchmark for evaluating vocal coaching judgments. This release contains expert annotations for 515 singing recordings: free-form coaching feedback, atomic diagnosis/correction claims, Top-3 issue labels, same-song triplet rankings, and segment-conditioned issue labels. Subsets: same_song / Dataset A: 207 Amazing Grace performances from DAMP-S-AG. Audio is not redistributed; use audio_filename to match the official release.… See the full description on the dataset page: https://huggingface.co/datasets/vocalcoachbench/vocalcoachbench-review.audioaudio-classification10K<n<100K0 likes87 downloads5mo agoHugging Face17isabeth /rgad-crosslingual-tts-10h RGAD Cross-Lingual TTS 10h This is a 10-hour cross-lingual TTS dataset for prompt-conditioned Chinese TTS fine-tuning. Format The dataset contains: train.jsonl dev.jsonl metadata.csv audio/prompts/*.wav audio/targets/*.wav Each JSONL row has this format: {"id":"sample_000001","prompt_wav":"audio/prompts/sample_000001.wav","target_wav":"audio/targets/sample_000001.wav","text":"中文目标文本。","prompt_language":"en-US","target_language":"zh-CN"… See the full description on the dataset page: https://huggingface.co/datasets/isabeth/rgad-crosslingual-tts-10h.audiotext-to-speech1K<n<10K1 likes80 downloads4mo agoHugging Face18umd-zhou-lab /AVQA-Audio-Rubrics AVQA Audio-Reasoning Rubrics Project Page | Paper | Code Audio-grounded, binary-evaluable evaluation rubrics for the full AVQA training set, generated for process-level reward modeling in audio reasoning RL (e.g. GRPO / RLHF with rubric-as-reward). Each training question is annotated with 5 rubrics, one per evaluation facet, that judge the quality of an audio-reasoning response — not just final answer correctness. The rubrics are designed to be scored Yes/No by an LLM judge that… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/AVQA-Audio-Rubrics.textaudio-classification10K<n<100K1 likes75 downloads2mo agoHugging Face19vhands /audio-reasoning-qa-post-public audio-reasoning-qa-post-public Question-answering and multi-task audio reasoning annotations across 15 public audio QA datasets. Spans general audio QA (Clotho-AQA, HeySQuAD), music reasoning (MU-LLaMA, MusicBench, LLARK-MTAT, Music-AVQA), speech-grounded QA (LibriSQA, GigaSpeech), and NVIDIA-aggregator skill subsets (TemporalQA, CountingQA, AudioSet-Speech-QA, GigaSpeech-Long-QA). Closes a substantial slice of the public audio-reasoning SFT gap (compare to NVIDIA AudioSkills-XL… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-reasoning-qa-post-public.textquestion-answering100K<n<1M0 likes68 downloads3mo agoHugging Face20Rakancorle1 /hans-10k Hans-10K · DPO recipe for the audio-visual Clever Hans DPO training data accompanying the paper When Vision Speaks for Sound. Like the original Clever Hans 🐎 — the horse that looked like he could do arithmetic but was actually reading his trainer's body language — video-capable MLLMs often look like they can hear: they answer audio questions by reading visual cues and never verifying the audio stream. Hans-10K is the 10,383-sample best-recipe preference-pair dataset that cures this… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-10k.audioaudio-classification10K<n<100K0 likes52 downloads4mo agoHugging Face21ASGPIPO /ryuu_lion_danmemoaudio1K<n<10K0 likes50 downloads18d agoHugging Face22Rakancorle1 /thud-eval THUD-Eval · audio-visual Clever Hans benchmark Evaluation benchmark accompanying the paper When Vision Speaks for Sound. This dataset probes the audio-visual Clever Hans effect — the tendency of video-capable MLLMs to appear to listen while really just reading visual cues. We test the same source clips under three audio interventions: Task Intervention What it tests sync audio temporally shifted (early / delay) Can the model detect a time offset? mute audio replaced… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/thud-eval.audion<1K0 likes44 downloads4mo agoHugging Face23electron-rare /mascarade-dsp-dataset Mascarade — DSP & Signal Processing Q&A ✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11) Per-sample Stack Exchange Electronics attribution recovered via the SE /search/advanced + /questions/{id} API search : 169 samples (~5.35 %) confirmed as Stack Exchange Electronics (CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution (URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60). 535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/mascarade-dsp-dataset.texttext-generation1K<n<10K0 likes30 downloads5mo agoHugging Face24Rakancorle1 /vggsync-3k VGGSync-3K · out-of-domain audio-visual sync benchmark Out-of-domain evaluation set used in the paper When Vision Speaks for Sound. Derived from VGGSoundSync, this 3,000-clip slice tests whether a video-capable MLLM can detect audio temporal offsets on everyday sound events outside the THUD in-domain training distribution. Each item is one VGGSound clip in one of three conditions: Condition Count gt_synced gt_direction gt_offset_sec Audio aligned (no shift) 1,000 true… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/vggsync-3k.audioaudio-classification1K<n<10K0 likes28 downloads4mo agoHugging Face25Rakancorle1 /hans-sft-4k Hans-SFT-4K · SFT recipe for the audio-visual Clever Hans Supervised fine-tuning (SFT) data accompanying the paper When Vision Speaks for Sound. Like the original Clever Hans — the horse that looked like he could do arithmetic but was actually reading his trainer's body language — video-capable MLLMs often look like they can hear: they answer audio questions by reading visual cues and never verifying the audio stream. Hans-SFT-4K is the 3,834-sample SFT mix that teaches models to… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-sft-4k.audioaudio-classification1K<n<10K1 likes24 downloads4mo agoHugging Face26rebellsport /my-first-moves my first moves • Reachy Mini Moves Community-contributed Marionette recordings captured on Reachy Mini. Files live under data/, each move ships as a JSON trajectory plus an optional WAV. Recorded with the Marionette web app. Reuse Cite this dataset as rebellsport/my-first-moves. Keep the reachy_mini_community_moves tag when sharing derivatives so the community can discover related sets. audioroboticsn<1K0 likes21 downloads1mo agoHugging Face27FBK-MT /RedVoxgated RedVox Multilingual red teaming dataset for audio and speech. This dataset corresponds to the test set presented in the paper "RedVox: Safety and Fairness Gaps in Speech Models Across Languages" (Savoldi, Papi et al., 2026) Dataset Structure The dataset is organized by language configuration: en/ - English language samples (1,359 entries) de/ - German language samples (519 entries) es/ - Spanish language samples (354 entries) fr/ - French language samples… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/RedVox.audio1K<n<10K1 likes19 downloads1mo agoHugging Face28RidheshBhati /Indic_New_dataset_TTS Indic TTS Dataset Hub (Mozilla) Validated audio–text pairs for multiple Indic languages from Mozilla Common Voice. Select the language from the Subset dropdown in the Dataset Viewer. Columns audio: WAV audio clip (16kHz) text: transcription duration: length in seconds speaking_rate: characters per second audio10K<n<100K0 likes12 downloads7mo agoHugging Face29tfrere /reachy-mini-onboarding-movesaudion<1K0 likes10 downloads3mo agoHugging Face30Reza2kn /homorich-negara-gooya-grapheme-regen-audioaudion<1K0 likes10 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.