datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VocalBench
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
This is the official release of VocalBench
Citation
If you find our work helpful, please cite our paper:
@article{liu2025vocalbench,
title={VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models},
author={Liu, Heyang and Wang, Yuhao and Cheng, Ziyang and Wu, Ronghua and Gu, Qunshan and Wang, Yanfeng and Wang, Yu},
journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/VocalNet/VocalBench.synthetic_vocal_burstsThis repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository.
https://huggingface.co/datasets/sleeping-ai/Vocal-burst
We captioned them using Gemini Flash Audio 2.0. This dataset contains, this dataset contains ~ 365,000 vocal bursts from all kinds of categories.
It might be helpful for pre-training audio text foundation models to generate and understand all kinds of nuances in vocal bursts.
vocalgrad
VocalGrad
VocalGrad is an audio benchmark for evaluating whether a model can detect the
direction of gradual perceptual change in speech. This public release contains
the test split only.
Each example contains one audio clip and one target attribute. The task is to
answer whether that attribute increases or decreases over time.
Task
Given an audio clip and an attribute name, predict one of two labels:
increase
decrease
The ground-truth label is derived from the metadata… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-user-592888/vocalgrad.Designed-Vocalizations-Dataset
Designed Vocalizations Dataset
Paper · Demo & audio samples
The Designed Vocalizations Dataset supports voice conversion for designed vocalizations
— monster growls, robotic voices, and other sound-designed timbres — an area left
underexplored by benchmarks that focus on natural human speech. It curates diverse raw vocal
sources (speech and animal / non-linguistic sounds) and applies professional vocal-effects
processing to produce corresponding effect-modified variants. A… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/Designed-Vocalizations-Dataset.vocalvocal_imitation_synth
Dataset Card for "vocal_imitation_synth"
More Information needed
vocalset_synthtext_interference_vocalsoundvocalset_synth
Dataset Card for "vocalset_synth"
More Information needed
vocal-affect-bench
VocalAffectBench
VocalAffectBench is a test-only benchmark for evaluating whether AI audio models can identify expressed vocal emotion from raw audio.
Paper: VocalAffectBench: Evaluating Vocal Emotion Recognition in AI Audio Models
The benchmark targets the expressed emotion — what the speaker conveys through vocal tone, prosody, pace, intensity, and pauses — not inferred internal state.
Contents
280 human-recorded English WAV clips, totalling 2.32 hours.
7… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/vocal-affect-bench.vocalsoundvocal-burst-classification-v2
Vocal Burst Classification V2 — laion/vocal-burst-classification-v2
The V2 training corpus for the Vocal Burst Classifier V2:
a single-label dataset over an 83-class vocal-burst taxonomy (82 non-speech human vocalizations
no_burst, index 82), shipped as precomputed VoiceCLAP-commercial embeddings plus the raw
vocal-bursts-clean audio.
The vocal-burst clips were generated with various synthetic text-to-audio models such as DramaBox,
then annotated and filtered as described… See the full description on the dataset page: https://huggingface.co/datasets/laion/vocal-burst-classification-v2.VoiceAssistant-430K-vocalnet
VoiceAssistant-430K-vocalnet
This dataset supports the reproduction of VocalNet.
Data Construction
Data Source: We used the VoiceAssistant-400K from Mini-Omni, which contains about 470K instances.
Data Filtering: We removed samples with excessively long data. The resulted corpus contains 430K instances.
Response Speech: We perform speech synthesis using CosyVoice to generate the response speech.
Response Token: We generate the speech token using CosyVoice2.… See the full description on the dataset page: https://huggingface.co/datasets/VocalNet/VoiceAssistant-430K-vocalnet.vocal-bursts
Vocal Bursts
A curated collection of 28,564 non-speech vocal burst audio samples across 18 categories.
Categories
Category
Samples
Breath
1,690
Cough
4,248
Crying
2,020
Laughter
4,797
Lip Popping
54
Lip Smacking
45
Moan
45
Nose Blowing
48
Pant
44
Scream
714
Sigh
3,546
Sneeze
3,813
Sniff
3,504
Teeth Chattering
46
Teeth Grinding
43
Throat Clearing
3,549
Tongue Clicking
47
Yawn
311
Format
Audio: FLAC
Metadata:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/vocal-bursts.vocal_bursts_taxonomy_100_clean_wdsvocal-money-codeswitch-asr-benchmark
Vocal Money — Yoruba–English Code-Switched ASR Benchmark
A 210-clip evaluation subset used to benchmark five speech recognition systems on naturally
code-switched Yoruba–English speech, together with the reference transcriptions and the output of
every system on every clip, so that the published results can be recomputed or contradicted.
Produced for the MLC (Africa) × Intron Agentic Voice AI Challenge, Deep Learning Indaba 2026.
Team Vocal Money — Hospice Hounfodji, Mohamed… See the full description on the dataset page: https://huggingface.co/datasets/Kimyayd/vocal-money-codeswitch-asr-benchmark.vocal-bursts-taxonomy-DACVAE
Vocal Bursts Taxonomy — DACVAE + MaestroClap Embeddings & Scores
Processed version of with DACVAE latents, MaestroClap embeddings, derived attribute/quality/speaker scores, and Gemini-verified labels.
Overview
Metric
Value
Total samples
16,175
Categories
82
Genders
male, female
Female samples
8,097
Male samples
8,078
Gemini Label Verification
Every sample was sent to Gemini 3.1 Flash Lite for two independent tasks:
Match scoring:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/vocal-bursts-taxonomy-DACVAE.vocal-burst-annotation-asr-tuning-dataset
Vocal Burst Annotation ASR Tuning Dataset
A synthetic 500,000-sample multilingual dataset for training ASR models with inline vocal burst captioning, speaker diarization, and sentence-level timestamps. Each sample is approximately 1 minute of audio containing speech segments interleaved with vocal bursts (laughs, sighs, coughs, etc.), annotated with precise timing information.
Example Transcript
[nasalized, affirmative hum, steady pitch, moderate intensity]… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/vocal-burst-annotation-asr-tuning-dataset.vocal_imitation_extract_unit
Dataset Card for "vocal_imitation_extract_unit"
More Information needed
vocalsound
This dataset only contains test data, which is integrated into UltraEval-Audio(https://github.com/OpenBMB/UltraEval-Audio) framework.
python audio_evals/main.py --dataset vocalsound --model gpt4o_audio
🚀超凡体验,尽在UltraEval-Audio🚀
UltraEval-Audio——全球首个同时支持语音理解和语音生成评估的开源框架,专为语音大模型评估打造,集合了34项权威Benchmark,覆盖语音、声音、医疗及音乐四大领域,支持十种语言,涵盖十二类任务。选择UltraEval-Audio,您将体验到前所未有的便捷与高效:
一键式基准管理 📥:告别繁琐的手动下载与数据处理,UltraEval-Audio为您自动化完成这一切,轻松获取所需基准测试数据。
内置评估利器… See the full description on the dataset page: https://huggingface.co/datasets/TwinkStart/vocalsound.vocalset_extract_unit
Dataset Card for "vocalset_extract_unit"
More Information needed
vocalcoachbench-review
VocalCoachBench
VocalCoachBench is a singing-audio benchmark for evaluating vocal coaching
judgments. This release contains expert annotations for 515 singing recordings:
free-form coaching feedback, atomic diagnosis/correction claims, Top-3 issue
labels, same-song triplet rankings, and segment-conditioned issue labels.
Subsets:
same_song / Dataset A: 207 Amazing Grace performances from DAMP-S-AG.
Audio is not redistributed; use audio_filename to match the official release.… See the full description on the dataset page: https://huggingface.co/datasets/vocalcoachbench/vocalcoachbench-review.vocal-score-synthetic-smoke-test
Vocal Score Synthetic Smoke Test
A tiny, fully synthetic fixture for smoke-testing monophonic vocal-to-notation pipelines. It contains one ten-second “ah”-like synthesized rendition of the public-domain melody commonly known as “Twinkle, Twinkle, Little Star,” its scripted note sequence, and observed exports from my Vocal Score research pipeline.
No human voice, copyrighted recording, voice embedding, lyric, or personal data is included.
Files
twinkle.wav — 22.05… See the full description on the dataset page: https://huggingface.co/datasets/mattwinwood/vocal-score-synthetic-smoke-test.vocalno
Tobias Chinese TTS Dataset
这是一个中文文本转语音(TTS)数据集,包含约997个高质量的中文音频-文本对。
数据集信息
语言: 中文 (Chinese)
任务: 文本转语音 (Text-to-Speech)
样本数量: ~997个音频-文本对
音频格式: WAV, 16kHz采样率
许可证: MIT
发言人: 单一发言人
数据集结构
from datasets import load_dataset
# 加载完整数据集
dataset = load_dataset("your_username/tobias-tts-chinese")
# 只加载训练集
train_dataset = load_dataset("your_username/tobias-tts-chinese", split="train")
# 只加载验证集
validation_dataset = load_dataset("your_username/tobias-tts-chinese"… See the full description on the dataset page: https://huggingface.co/datasets/t0bi4s/vocalno.Audio-Hallucination_Object-Existence_AudioCaps-ESC50-VocalSound
Dataset Card for "Audio-Hallucination_Object-Existence_AudioCaps-ESC50-VocalSound"
More Information needed
speech2speech_vocalnetCS3-Bench
CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
Github Repo: https://github.com/SJTU-OmniAgent/CS3-Bench
CS3-Bench is accepted at ICASSP 2026 conference.
This repository hosts CS3-Bench, a Code-Switching Speech-to-Speech Benchmark dataset, as presented in the paper CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching.
The benchmark is designed to evaluate and improve the language alignment… See the full description on the dataset page: https://huggingface.co/datasets/VocalNet/CS3-Bench.wan_gen_videos_HunyuanVideo_Foley_no_vocal_sound_captionedvocalset_unitvocal-bursts
Vocal Bursts
A curated collection of 28,564 non-speech vocal burst audio samples across 18 categories.
Categories
Category
Samples
Breath
1,690
Cough
4,248
Crying
2,020
Laughter
4,797
Lip Popping
54
Lip Smacking
45
Moan
45
Nose Blowing
48
Pant
44
Scream
714
Sigh
3,546
Sneeze
3,813
Sniff
3,504
Teeth Chattering
46
Teeth Grinding
43
Throat Clearing
3,549
Tongue Clicking
47
Yawn
311
Format
Audio: FLAC… See the full description on the dataset page: https://huggingface.co/datasets/0x3/vocal-bursts.
