datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Audio2Tool
Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use
Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1
1 Rivian & Volkswagen Technologies · ∗ equal contribution · ∗∗ corresponding author · † equal contribution
📄 Project page / demo: https://audio2tool.github.io/
📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool
✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.MixEval-X
🚀 Project Page | 📜 arXiv | 👨💻 Github | 🏆 Leaderboard | 📝 blog | 🤗 HF Paper | 𝕏 Twitter
MixEval-X encompasses eight input-output modality combinations and can be further extended. Its data points reflect real-world task distributions. The last grid presents the scores of frontier organizations’ flagship models on MixEval-X, normalized to a 0-100 scale, with MMG tasks using win rates instead of Elo. Section C of the paper presents example data samples and model responses.… See the full description on the dataset page: https://huggingface.co/datasets/MixEval/MixEval-X.jam-actions-v1
jam-actions-v1
Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) ·
Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE ·
Source repo: mcp-tool-shop-org/ai-jam-sessions
The successor to jam-actions-v0.
Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from
what the tools return — and it exists in its current shape because, seven training runs in a row,
the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.jam-actions-acoustic-v0
Dataset Card for jam-actions-acoustic-v0
Version: 1.0.2
Published at mcp-tool-shop/jam-actions-acoustic-v0. No DOI.
Summary
108 constructible gold records of grounded MCP tool use over monophonic audio analysis. Each record pairs a 4-note right-hand reduction of a public-domain library phrase with a seeded synthetic take and a gold verdict (match, pitch fail/warn, timing fail/pass, missed, extra, in-tune vibrato, or nothing-to-grade silence).
This is not a musical… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-acoustic-v0.audio-event-classification-post-public
audio-event-classification-post-public
Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.appointment-bench
Appointment Bench
25-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a dental office receptionist handling appointment scheduling.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a dental office receptionist scheduling appointments for two patients with confusable names (Daniel and Danielle Nolan)… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/appointment-bench.grocery-bench
Grocery Bench
30-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a grocery ordering assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a grocery ordering assistant helping a customer build, modify, and finalize an order. The conversation is designed around 15 difficulty enhancements that… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/grocery-bench.product-bench
Product Bench
31-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a laptop comparison shopping assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a laptop comparison shopping assistant helping a customer evaluate, compare, and order laptops. The conversation features multi-intent turns… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/product-bench.vocalcoachbench-review
VocalCoachBench
VocalCoachBench is a singing-audio benchmark for evaluating vocal coaching
judgments. This release contains expert annotations for 515 singing recordings:
free-form coaching feedback, atomic diagnosis/correction claims, Top-3 issue
labels, same-song triplet rankings, and segment-conditioned issue labels.
Subsets:
same_song / Dataset A: 207 Amazing Grace performances from DAMP-S-AG.
Audio is not redistributed; use audio_filename to match the official release.… See the full description on the dataset page: https://huggingface.co/datasets/vocalcoachbench/vocalcoachbench-review.event-bench
Event Bench
29-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as an event planning assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as an event planning assistant managing venue bookings, catering, and guest logistics. The conversation features cascading changes — a venue switch triggers catering… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/event-bench.conversation-bench
Conversation Bench
75-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a conference assistant for the AI Engineer World's Fair.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a conference assistant for the AI Engineer World's Fair, handling session registration, schedule queries, speaker lookups, and… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/conversation-bench.audio-music-mir-post-public
audio-music-mir-post-public
Music information retrieval and tagging annotations: genre (FMA), instrument family (NSynth × 3, Medley-solos-DB), social tags (MagnaTagATune via LLARK), and large-scale Music4All metadata. Foundation for music understanding heads in audio LLMs.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py after fetching to rewrite the JSONL audio_path fields with absolute local… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-music-mir-post-public.assistant-bench
Assistant Bench
31-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a personal assistant handling flights, email, calendar, and reminders.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a personal assistant managing flight bookings, email composition, calendar events, and reminders. Turns include dual… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/assistant-bench.audio-reasoning-qa-post-public
audio-reasoning-qa-post-public
Question-answering and multi-task audio reasoning annotations across 15 public audio QA datasets. Spans general audio QA (Clotho-AQA, HeySQuAD), music reasoning (MU-LLaMA, MusicBench, LLARK-MTAT, Music-AVQA), speech-grounded QA (LibriSQA, GigaSpeech), and NVIDIA-aggregator skill subsets (TemporalQA, CountingQA, AudioSet-Speech-QA, GigaSpeech-Long-QA). Closes a substantial slice of the public audio-reasoning SFT gap (compare to NVIDIA AudioSkills-XL… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-reasoning-qa-post-public.ONOTE
ONOTE: Omnimodal Notation Objective Topology Examination
ONOTE is a large-scale, omnimodal benchmark designed to evaluate music understanding across three major notation systems: Western Staff, Jianpu (Numbered Notation), and Guitar Tablature.
📂 Dataset Structure
The dataset is organized into two primary sub-directories based on the notation and instrument type:
1. pitch_Jianpu_dataset (Staff & Numbered Notation)
This sub-folder focuses on Western staff and… See the full description on the dataset page: https://huggingface.co/datasets/Weisiqing123/ONOTE.Gurbani-MahanKosh-Frontier-Corpus
ੴ Gurbani & Bhai Kahn Singh Nabha Mahan Kosh Frontier Corpus
☬ ਗੁਰਬਾਣੀ ਅਤੇ ਭਾਈ ਕਾਹਨ ਸਿੰਘ ਨਾਭਾ 'ਮਹਾਨ ਕੋਸ਼' ਪ੍ਰਮਾਣਿਕ ਡਾਟਾਸੈੱਟ
👨💻 Project Lead & Architecture
Curator & Developer: Gurpreet Singh Dhillon (Nam-toon Studio)
GitHub Profile: github.com/gurpreetsingh5523-source
Flagship Project: AMRIT Research OS (Autonomous Medical AI)
📖 Dataset Overview
An authoritative lexical dataset compiling authentic definitions… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Gurbani-MahanKosh-Frontier-Corpus.mascarade-dsp-dataset
Mascarade — DSP & Signal Processing Q&A
✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11)
Per-sample Stack Exchange Electronics attribution recovered via the SE
/search/advanced + /questions/{id} API search :
169 samples (~5.35 %) confirmed as Stack Exchange Electronics
(CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution
(URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60).
535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/mascarade-dsp-dataset.TheresaOnomatopoeia_Dataset🎧 Onomatopoeia Dataset (Audio → Manga Expression)
音声解析結果をもとに、日本語のオノマトペ(擬音語・擬態語)を生成するためのデータセットです。
本データセットは、音そのものではなく、音から推定された特徴・空間・情景を入力とする構造化データであり、
漫画的な表現生成を目的としたマルチモーダルデータです。
📌 Dataset Summary
本データセットは以下のパイプラインから生成されています:
Audio
↓
Audio Features (04_features.json)
↓
Audio Events (05_audio_events.json)
↓
Space Judgement (06_space_judgement.json)
↓
Scene Interpretation (07_scene_interpretation.json)
↓
Onomatopoeia (08_onomatopoeia.json)
👉 音 → 空間 → 情景 → オノマトペ
という段階的生成構造を持ちます。
📊… See the full description on the dataset page: https://huggingface.co/datasets/yadorigi/Onomatopoeia_Dataset.EdgeMMEval
EdgeMMEval
Minimal multimodal evaluation dataset for on-device inference testing.
Covers functional correctness, accuracy, latency stress, and memory
pressure across image, audio, text, multi-turn, combination, structured
output, and tool-calling cases.
Dataset summary
The test split is defined in data/test/metadata.jsonl (200 rows). Each
row has a test_id (for example IMG-001, STO-020) and a modality.
Modality
Samples
Focus
Image
34
VQA, OCR, description… See the full description on the dataset page: https://huggingface.co/datasets/CortexSwarm/EdgeMMEval.txt2vst
txt2vst Dataset
10,000 natural-language-to-VST-spec pairs for music plugin generation.
Dataset Description
Each sample maps a natural language description of a VST instrument to a structured spec.json that defines the complete plugin architecture.
Fields
prompt: Natural language description (e.g., "drum machine with kick snare and acid bass, punchy mastering")
completion: Compact JSON spec defining the plugin (voices, FX, theme, mastering chain)… See the full description on the dataset page: https://huggingface.co/datasets/fabriziosalmi/txt2vst.ONOTE
ONOTE: Omnimodal Notation Objective Topology Examination
ONOTE is a large-scale, omnimodal benchmark designed to evaluate music understanding across three major notation systems: Western Staff, Jianpu (Numbered Notation), and Guitar Tablature.
📂 Dataset Structure
The dataset is organized into two primary sub-directories based on the notation and instrument type:
1. pitch_Jianpu_dataset (Staff & Numbered Notation)
This sub-folder focuses on Western staff and… See the full description on the dataset page: https://huggingface.co/datasets/chengxin666/ONOTE.synthetic-patient-dr-data
Synthetic Patient DR Data
Synthetic doctor-patient consultation dataset with structured clinical outputs and optional full-consultation audio.
Dataset Summary
This dataset was generated for research and prototyping in:
clinical dialogue generation
structured clinical extraction
text-to-audio workflows
conversational healthcare modeling
All consultations are synthetic and should not be treated as real clinical encounters.
Export Metadata
Mode: audio
Repo… See the full description on the dataset page: https://huggingface.co/datasets/TumeloKonaite/synthetic-patient-dr-data.mascarade-dsp-dataset
Mascarade — DSP & Signal Processing Q&A
✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11)
Per-sample Stack Exchange Electronics attribution recovered via the SE
/search/advanced + /questions/{id} API search :
169 samples (~5.35 %) confirmed as Stack Exchange Electronics
(CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution
(URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60).
535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-dsp-dataset.
