vocal
Datasets
All datasets matching “vocal”Barkopedia-Dog-Vocal-Detection
🐾 Dog Vocal Detection
This dataset is curated from internet videos to support research in dog vocalization detection using both weak and strong supervision.
It contains approximately 7,500 seconds of strongly labeled training audio
Over 9,000 seconds of weakly labeled clips sourced from AudioSet are included.
The dataset also provides 24 hours of unlabeled audio clips from our own collection.
To simulate realistic conditions, some clips feature dogs present without barking… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia-Dog-Vocal-Detection.VocalBench
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
This is the official release of VocalBench
Citation
If you find our work helpful, please cite our paper:
@article{liu2025vocalbench,
title={VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models},
author={Liu, Heyang and Wang, Yuhao and Cheng, Ziyang and Wu, Ronghua and Gu, Qunshan and Wang, Yanfeng and Wang, Yu},
journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/VocalNet/VocalBench.synthetic_vocal_burstsThis repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository.
https://huggingface.co/datasets/sleeping-ai/Vocal-burst
We captioned them using Gemini Flash Audio 2.0. This dataset contains, this dataset contains ~ 365,000 vocal bursts from all kinds of categories.
It might be helpful for pre-training audio text foundation models to generate and understand all kinds of nuances in vocal bursts.
vocalgrad
VocalGrad
VocalGrad is an audio benchmark for evaluating whether a model can detect the
direction of gradual perceptual change in speech. This public release contains
the test split only.
Each example contains one audio clip and one target attribute. The task is to
answer whether that attribute increases or decreases over time.
Task
Given an audio clip and an attribute name, predict one of two labels:
increase
decrease
The ground-truth label is derived from the metadata… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-user-592888/vocalgrad.scotus-vocal-data
Mimir: Supreme Court voice and vote data
Mimir studies whether measurements of justices' speech during oral argument help predict
their votes. Start with the evaluated model and its model card:
Mimir vote predictor:
the selected model, portable inference, final evaluation, coverage and limitations.
Complete supporting model data:
fitting data, separate final inputs and targets, candidate models, predictions, acoustic
measurements, new word and diarization outputs, source… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/scotus-vocal-data.Designed-Vocalizations-Dataset
Designed Vocalizations Dataset
Paper · Demo & audio samples
The Designed Vocalizations Dataset supports voice conversion for designed vocalizations
— monster growls, robotic voices, and other sound-designed timbres — an area left
underexplored by benchmarks that focus on natural human speech. It curates diverse raw vocal
sources (speech and animal / non-linguistic sounds) and applies professional vocal-effects
processing to produce corresponding effect-modified variants. A… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/Designed-Vocalizations-Dataset.
