datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic_vocal_burstsThis repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository.
https://huggingface.co/datasets/sleeping-ai/Vocal-burst
We captioned them using Gemini Flash Audio 2.0. This dataset contains, this dataset contains ~ 365,000 vocal bursts from all kinds of categories.
It might be helpful for pre-training audio text foundation models to generate and understand all kinds of nuances in vocal bursts.
more-synthetic-vocalbursts-raw
More Synthetic Vocal Bursts (Raw)
Synthetic vocal burst audio samples generated from a taxonomy of 202 vocal burst types across multiple text-to-audio and TTS models. Each sample is a short (3–10 second) non-speech vocalization — laughs, cries, gasps, sighs, growls, etc. — generated from text prompts describing the burst type, gender, and age group.
Models Used
Model
Type
Samples
Sample Rate
Notes
DramaBox (ResembleAI/Dramabox)
TTS DiT
2000
44.1 kHz… See the full description on the dataset page: https://huggingface.co/datasets/laion/more-synthetic-vocalbursts-raw.improved-synthetic-vocal-burtsimproved_synthetic_vocal_burtssynthetic-vocal-burstsimproved_synthetic_vocal_burtssynthetic-utterancessynthetic-dataset-v2
