datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Audio2Tool
Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use
Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1
1 Rivian & Volkswagen Technologies · ∗ equal contribution · ∗∗ corresponding author · † equal contribution
📄 Project page / demo: https://audio2tool.github.io/
📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool
✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.gdpval_preference_rubricsmicroduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.XModBenchXModBench
Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
🎉 Accepted at ICLR 2026
What is XModBench?
XModBench is the first tri-modal (audio / vision / text) multiple-choice
QA benchmark explicitly designed to measure cross-modal consistency — does
an omni-language model give the same correct answer when the same semantic
content is presented in different modalities?
Each item is a 4-choice question with a <context>… See the full description on the dataset page: https://huggingface.co/datasets/RyanWW/XModBench.MMAG
MMAG: A Multi‑Control Mixed Audio Generation Benchmark
MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment.
Dataset Structure
The dataset is organized into three subsets, each targeting a specific… See the full description on the dataset page: https://huggingface.co/datasets/rookie9/MMAG.asr-reference-set-eval-temp
Temporary ASR evaluation audio
Temporary public audio files used for hosted ASR evaluation.
RAIL
RAIL Audio Benchmark
This folder is generated for direct Hugging Face dataset upload. Each row uses relative audio paths rooted at this repository folder.
NeurIPS / Croissant Metadata
metadata.json is a Croissant-style metadata file with core fields and minimal RAI fields.
metadata.json includes the Hugging Face dataset URL, CC-BY-4.0 license URL, checksums, and RAI fields.
build_summary.json contains build counts and skipped-source diagnostics.
Schema
id:… See the full description on the dataset page: https://huggingface.co/datasets/AnnoymousNeurLPSsubmit/RAIL.Seamless_Dummy_Dataset_Fixed_3
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
SlideASR-BenchCS50-rawmy-voxtral-datasetResource2Skill
Resource2Skill — Skill Library
Executable skill libraries for the Resource2Skill
runtime: reusable, structured skills that a software agent browses, inspects,
and composes to operate real tools (Web, PowerPoint, Excel, Blender, and
REAPER-style audio) and produce artifacts.
This dataset is the skill data half of the project; the runnable runtime,
MCP servers, and CLI live in the code repository.
Contents
skills_wiki/ Structured wiki entries used for runtime… See the full description on the dataset page: https://huggingface.co/datasets/YijiaFan/Resource2Skill.ableton-racks-labeled
Dataset Card for Ableton Live XML Presets
Dataset Summary
This dataset consists of Ableton Live XML presets paired with synthetic, multi-perspective natural language descriptions and Chain-of-Thought (CoT) reasoning. It is designed for fine-tuning Large Language Models (LLMs) to perform text-to-preset generation, sound design analysis, and audio parameter reasoning.
Each entry provides three distinct prompt styles (layman, technical, and emotional) alongside a… See the full description on the dataset page: https://huggingface.co/datasets/sixf0ur/ableton-racks-labeled.ContextTTS_dataset
ContextTTS Evaluation Dataset
This is the official evaluation dataset for the paper "[ContextTTS Eval: A Benchmark for Evaluating Long-Form
Contextual Expressive Text-to-Speech]". It is designed to evaluate the performance of multi-modal speech synthesis, specifically focusing on context-aware prosody and timbre consistency in Chinese conversations and audiobooks.
Dataset Summary
The dataset consists of high-quality Chinese audio-text pairs, organized into three distinct… See the full description on the dataset page: https://huggingface.co/datasets/rodenhhh/ContextTTS_dataset.kencorpus_sw_culture
KenCorpus Swahili Culture Subset
A filtered subset of Kencorpus/KenCorpus_audio,
containing only rows where language=Swahili and genre=Culture (37 clips).
Audio files are in audio/, indexed by kencorpus_sw_culture.jsonl with path and duration fields,
following the layout of kyutai/DailyTalkContiguous.
vocalcoachbench-review
VocalCoachBench
VocalCoachBench is a singing-audio benchmark for evaluating vocal coaching
judgments. This release contains expert annotations for 515 singing recordings:
free-form coaching feedback, atomic diagnosis/correction claims, Top-3 issue
labels, same-song triplet rankings, and segment-conditioned issue labels.
Subsets:
same_song / Dataset A: 207 Amazing Grace performances from DAMP-S-AG.
Audio is not redistributed; use audio_filename to match the official release.… See the full description on the dataset page: https://huggingface.co/datasets/vocalcoachbench/vocalcoachbench-review.rgad-crosslingual-tts-10h
RGAD Cross-Lingual TTS 10h
This is a 10-hour cross-lingual TTS dataset for prompt-conditioned Chinese TTS fine-tuning.
Format
The dataset contains:
train.jsonl
dev.jsonl
metadata.csv
audio/prompts/*.wav
audio/targets/*.wav
Each JSONL row has this format:
{"id":"sample_000001","prompt_wav":"audio/prompts/sample_000001.wav","target_wav":"audio/targets/sample_000001.wav","text":"中文目标文本。","prompt_language":"en-US","target_language":"zh-CN"… See the full description on the dataset page: https://huggingface.co/datasets/isabeth/rgad-crosslingual-tts-10h.AVQA-Audio-Rubrics
AVQA Audio-Reasoning Rubrics
Project Page | Paper | Code
Audio-grounded, binary-evaluable evaluation rubrics for the full
AVQA training set, generated for
process-level reward modeling in audio reasoning RL (e.g. GRPO / RLHF with
rubric-as-reward).
Each training question is annotated with 5 rubrics, one per evaluation
facet, that judge the quality of an audio-reasoning response — not just final
answer correctness. The rubrics are designed to be scored Yes/No by an
LLM judge that… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/AVQA-Audio-Rubrics.audio-reasoning-qa-post-public
audio-reasoning-qa-post-public
Question-answering and multi-task audio reasoning annotations across 15 public audio QA datasets. Spans general audio QA (Clotho-AQA, HeySQuAD), music reasoning (MU-LLaMA, MusicBench, LLARK-MTAT, Music-AVQA), speech-grounded QA (LibriSQA, GigaSpeech), and NVIDIA-aggregator skill subsets (TemporalQA, CountingQA, AudioSet-Speech-QA, GigaSpeech-Long-QA). Closes a substantial slice of the public audio-reasoning SFT gap (compare to NVIDIA AudioSkills-XL… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-reasoning-qa-post-public.hans-10k
Hans-10K · DPO recipe for the audio-visual Clever Hans
DPO training data accompanying the paper
When Vision Speaks for Sound.
Like the original Clever Hans 🐎 —
the horse that looked like he could do arithmetic but was actually reading
his trainer's body language — video-capable MLLMs often look like they
can hear: they answer audio questions by reading visual cues and never
verifying the audio stream.
Hans-10K is the 10,383-sample best-recipe preference-pair dataset
that cures this… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-10k.ryuu_lion_danmemothud-eval
THUD-Eval · audio-visual Clever Hans benchmark
Evaluation benchmark accompanying the paper
When Vision Speaks for Sound.
This dataset probes the audio-visual Clever Hans effect — the tendency
of video-capable MLLMs to appear to listen while really just reading
visual cues. We test the same source clips under three audio
interventions:
Task
Intervention
What it tests
sync
audio temporally shifted (early / delay)
Can the model detect a time offset?
mute
audio replaced… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/thud-eval.mascarade-dsp-dataset
Mascarade — DSP & Signal Processing Q&A
✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11)
Per-sample Stack Exchange Electronics attribution recovered via the SE
/search/advanced + /questions/{id} API search :
169 samples (~5.35 %) confirmed as Stack Exchange Electronics
(CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution
(URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60).
535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/mascarade-dsp-dataset.vggsync-3k
VGGSync-3K · out-of-domain audio-visual sync benchmark
Out-of-domain evaluation set used in the paper
When Vision Speaks for Sound.
Derived from VGGSoundSync,
this 3,000-clip slice tests whether a video-capable MLLM can detect
audio temporal offsets on everyday sound events outside the THUD
in-domain training distribution.
Each item is one VGGSound clip in one of three conditions:
Condition
Count
gt_synced
gt_direction
gt_offset_sec
Audio aligned (no shift)
1,000
true… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/vggsync-3k.hans-sft-4k
Hans-SFT-4K · SFT recipe for the audio-visual Clever Hans
Supervised fine-tuning (SFT) data accompanying the paper
When Vision Speaks for Sound.
Like the original Clever Hans —
the horse that looked like he could do arithmetic but was actually reading
his trainer's body language — video-capable MLLMs often look like they
can hear: they answer audio questions by reading visual cues and never
verifying the audio stream.
Hans-SFT-4K is the 3,834-sample SFT mix that teaches models to… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-sft-4k.my-first-moves
my first moves • Reachy Mini Moves
Community-contributed Marionette recordings captured on Reachy Mini.
Files live under data/, each move ships as a JSON trajectory plus an optional WAV.
Recorded with the Marionette web app.
Reuse
Cite this dataset as rebellsport/my-first-moves.
Keep the reachy_mini_community_moves tag when sharing derivatives so the community can discover related sets.
RedVox
RedVox
Multilingual red teaming dataset for audio and speech.
This dataset corresponds to the test set presented in the paper
"RedVox: Safety and Fairness Gaps in Speech Models Across Languages" (Savoldi, Papi et al., 2026)
Dataset Structure
The dataset is organized by language configuration:
en/ - English language samples (1,359 entries)
de/ - German language samples (519 entries)
es/ - Spanish language samples (354 entries)
fr/ - French language samples… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/RedVox.Indic_New_dataset_TTS
Indic TTS Dataset Hub (Mozilla)
Validated audio–text pairs for multiple Indic languages from Mozilla Common Voice.
Select the language from the Subset dropdown in the Dataset Viewer.
Columns
audio: WAV audio clip (16kHz)
text: transcription
duration: length in seconds
speaking_rate: characters per second
reachy-mini-onboarding-moveshomorich-negara-gooya-grapheme-regen-audio
