datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Audio2Tool
Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use
Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1
1 Rivian & Volkswagen Technologies · ∗ equal contribution · ∗∗ corresponding author · † equal contribution
📄 Project page / demo: https://audio2tool.github.io/
📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool
✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text
Dataset Overview
A collection of 27 domains (“topics”) and 3100 question-answer pair.
Each topic comes with average 117 QA pairs.Every QA entry comes with:
references: one or more source files the answer is extracted from
time with each reference comes the starting and ending time the answer is extracted from the reference
video_files: the video files where the answer can be found
(future) video title & description from metadata.csv
File structure
You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.SwitchLingua_audio
Dataset Card for SwitchLingua_text
🚀 News
[19/09/2025] SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset is accepted by NeurIPS 2025!
[30/05/2024] The manuscript can be found on arXiv.
Dataset Summary
SwitchLingua is a comprehensive multilingual and multicultural code-switching dataset designed to advance research in automatic speech recognition, natural language processing, and conversational AI. The… See the full description on the dataset page: https://huggingface.co/datasets/Shelton1013/SwitchLingua_audio.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.pashto-audio-wav2vecaudio-event-classification-post-public
audio-event-classification-post-public
Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.Ficbook-Audio-Instruct-10K
Ficbook Audio Instruct 10K
Synthetic audio instruction dataset for training Russian audio-language models.
Contains ~10K samples of fiction text voiced with OpenAI TTS and paired with diverse instruction tasks.
Dataset Description
This dataset was created for training and evaluating audio-language models on Russian fiction content.
Each sample contains:
Audio: Fiction text voiced using OpenAI's gpt-4o-mini-tts model
Text: Original text from ficbook stories
Question:… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Ficbook-Audio-Instruct-10K.audio-music-mir-post-public
audio-music-mir-post-public
Music information retrieval and tagging annotations: genre (FMA), instrument family (NSynth × 3, Medley-solos-DB), social tags (MagnaTagATune via LLARK), and large-scale Music4All metadata. Foundation for music understanding heads in audio LLMs.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py after fetching to rewrite the JSONL audio_path fields with absolute local… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-music-mir-post-public.audio-reasoning-qa-post-public
audio-reasoning-qa-post-public
Question-answering and multi-task audio reasoning annotations across 15 public audio QA datasets. Spans general audio QA (Clotho-AQA, HeySQuAD), music reasoning (MU-LLaMA, MusicBench, LLARK-MTAT, Music-AVQA), speech-grounded QA (LibriSQA, GigaSpeech), and NVIDIA-aggregator skill subsets (TemporalQA, CountingQA, AudioSet-Speech-QA, GigaSpeech-Long-QA). Closes a substantial slice of the public audio-reasoning SFT gap (compare to NVIDIA AudioSkills-XL… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-reasoning-qa-post-public.audio-function-calling
Audio Function Calling Dataset
A synthetic multi-turn conversation dataset for audio-based tool/function calling.
Each sample contains a system prompt with tool definitions, alternating user and assistant turns,
where user turns are designed for audio (natural spoken language with filler words) and assistant
turns may include tool calls.
Note: This is the first batch (137 audio samples, 153 text-only samples). We are actively improving the generation pipeline and will be adding… See the full description on the dataset page: https://huggingface.co/datasets/mlech26l/audio-function-calling.audio-alpacaaudio-function-calling
Audio Function Calling Dataset
A synthetic multi-turn conversation dataset for audio-based tool/function calling.
Each sample contains a system prompt with tool definitions, alternating user and assistant turns,
where user turns are designed for audio (natural spoken language with filler words) and assistant
turns may include tool calls.
Note: This is the first batch (137 audio samples, 153 text-only samples). We are actively improving the generation pipeline and will be adding… See the full description on the dataset page: https://huggingface.co/datasets/nandinireddy123/audio-function-calling.medreport_audio_204
MedReport - Audio Dataset
Dataset Description
This dataset contains medical report audio files with their transcriptions, formatted according to HuggingFace Audio Dataset specifications. It's suitable for training speech-to-text models and instruction-following models in the medical domain.
Dataset Structure
This dataset follows the official HuggingFace Audio Dataset format:
dataset/
└── train/
├── audio/
│ ├── 20240315143022.wav
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/wouk1805/medreport_audio_204.gene-speech-audio-instruct
speech-audio-instruct v3
Gate-passed instruction data for speech-audio — published when 50 fresh examples cleared the quality bar
Kind: synthetic
Domain: speech-audio
Records: 139
Created: 2026-06-25T15:25:34+00:00
SHA-256: c388150c1632b511928f0569229ed89fa9d28953fe4c9ae44811f8ed4eaf45d2
Pipeline: v2.0.0
Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": "llama", "min_judge": 0.7}
Generated by: Qwen3-4B-Instruct-2507-Q4_K_M.gguf (backend: llama)… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-speech-audio-instruct.audiodvdluanblazing-audio-slm-v7-4-dataset-dev
Blazing Audio SLM V7.4 development dataset
This is the exact joint-replay optimizer dataset used for the V7.4 development checkpoint. It is
not a claim that the resulting model passed release: V7.4 failed the minimum-family,
general-adversarial, and strict measurement-deferral V1 gates. Manifest V2 remains sealed and is
not included.
Splits and composition
Split
Rows
Calculate
Explain
Abstain
Defer measurement
Train
3,923
2,621
582
360
360
Validation… See the full description on the dataset page: https://huggingface.co/datasets/audiuphile/blazing-audio-slm-v7-4-dataset-dev.tw-daily-dialogue-audio
Dataset Card for tw-daily-dialogue-audio
本資料集是一份臺灣日常情境的對話腳本(dialogue scripts)資料集,每筆樣本包含對話分類、主題、文字內容以及說話者輪廓/場景/天氣等情境 metadata。可作為文字→語音(TTS)合成、對話 ASR 評測之素材設計來源。共 29,337 筆樣本。
Dataset Details
Dataset Description
資料以對話腳本(純文字)為主,搭配豐富的情境 metadata:
category:對話類別(如:餐廳、醫療、家庭、商業、交通等)。
theme:該段對話的細部主題。
text:對話腳本本文(可包含多輪、多角色)。
meta:情境 metadata,包含:
profile:說話者輪廓
loc:場景/地點
weather:當下天氣
可用於下游語音/對話應用:將腳本送入 TTS 管線生成多角色音訊、評測 dialogue-aware ASR、訓練具備情境感知的對話模型。… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-daily-dialogue-audio.tdtu-vi-audio
