datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Audio2Tool
Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use
Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1
1 Rivian & Volkswagen Technologies · ∗ equal contribution · ∗∗ corresponding author · † equal contribution
📄 Project page / demo: https://audio2tool.github.io/
📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool
✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.SwitchLingua_audio
Dataset Card for SwitchLingua_text
🚀 News
[19/09/2025] SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset is accepted by NeurIPS 2025!
[30/05/2024] The manuscript can be found on arXiv.
Dataset Summary
SwitchLingua is a comprehensive multilingual and multicultural code-switching dataset designed to advance research in automatic speech recognition, natural language processing, and conversational AI. The… See the full description on the dataset page: https://huggingface.co/datasets/Shelton1013/SwitchLingua_audio.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.pashto-audio-wav2vecaudio-event-classification-post-public
audio-event-classification-post-public
Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.Ficbook-Audio-Instruct-10K
Ficbook Audio Instruct 10K
Synthetic audio instruction dataset for training Russian audio-language models.
Contains ~10K samples of fiction text voiced with OpenAI TTS and paired with diverse instruction tasks.
Dataset Description
This dataset was created for training and evaluating audio-language models on Russian fiction content.
Each sample contains:
Audio: Fiction text voiced using OpenAI's gpt-4o-mini-tts model
Text: Original text from ficbook stories
Question:… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Ficbook-Audio-Instruct-10K.audio-music-mir-post-public
audio-music-mir-post-public
Music information retrieval and tagging annotations: genre (FMA), instrument family (NSynth × 3, Medley-solos-DB), social tags (MagnaTagATune via LLARK), and large-scale Music4All metadata. Foundation for music understanding heads in audio LLMs.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py after fetching to rewrite the JSONL audio_path fields with absolute local… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-music-mir-post-public.audio-reasoning-qa-post-public
audio-reasoning-qa-post-public
Question-answering and multi-task audio reasoning annotations across 15 public audio QA datasets. Spans general audio QA (Clotho-AQA, HeySQuAD), music reasoning (MU-LLaMA, MusicBench, LLARK-MTAT, Music-AVQA), speech-grounded QA (LibriSQA, GigaSpeech), and NVIDIA-aggregator skill subsets (TemporalQA, CountingQA, AudioSet-Speech-QA, GigaSpeech-Long-QA). Closes a substantial slice of the public audio-reasoning SFT gap (compare to NVIDIA AudioSkills-XL… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-reasoning-qa-post-public.audio-function-calling
Audio Function Calling Dataset
A synthetic multi-turn conversation dataset for audio-based tool/function calling.
Each sample contains a system prompt with tool definitions, alternating user and assistant turns,
where user turns are designed for audio (natural spoken language with filler words) and assistant
turns may include tool calls.
Note: This is the first batch (137 audio samples, 153 text-only samples). We are actively improving the generation pipeline and will be adding… See the full description on the dataset page: https://huggingface.co/datasets/mlech26l/audio-function-calling.medreport_audio_204
MedReport - Audio Dataset
Dataset Description
This dataset contains medical report audio files with their transcriptions, formatted according to HuggingFace Audio Dataset specifications. It's suitable for training speech-to-text models and instruction-following models in the medical domain.
Dataset Structure
This dataset follows the official HuggingFace Audio Dataset format:
dataset/
└── train/
├── audio/
│ ├── 20240315143022.wav
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/wouk1805/medreport_audio_204.audio-function-calling
Audio Function Calling Dataset
A synthetic multi-turn conversation dataset for audio-based tool/function calling.
Each sample contains a system prompt with tool definitions, alternating user and assistant turns,
where user turns are designed for audio (natural spoken language with filler words) and assistant
turns may include tool calls.
Note: This is the first batch (137 audio samples, 153 text-only samples). We are actively improving the generation pipeline and will be adding… See the full description on the dataset page: https://huggingface.co/datasets/nandinireddy123/audio-function-calling.tw-daily-dialogue-audio
Dataset Card for tw-daily-dialogue-audio
本資料集是一份臺灣日常情境的對話腳本(dialogue scripts)資料集,每筆樣本包含對話分類、主題、文字內容以及說話者輪廓/場景/天氣等情境 metadata。可作為文字→語音(TTS)合成、對話 ASR 評測之素材設計來源。共 29,337 筆樣本。
Dataset Details
Dataset Description
資料以對話腳本(純文字)為主,搭配豐富的情境 metadata:
category:對話類別(如:餐廳、醫療、家庭、商業、交通等)。
theme:該段對話的細部主題。
text:對話腳本本文(可包含多輪、多角色)。
meta:情境 metadata,包含:
profile:說話者輪廓
loc:場景/地點
weather:當下天氣
可用於下游語音/對話應用:將腳本送入 TTS 管線生成多角色音訊、評測 dialogue-aware ASR、訓練具備情境感知的對話模型。… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-daily-dialogue-audio.tdtu-vi-audio
