datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic_vocal_burstsThis repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository.
https://huggingface.co/datasets/sleeping-ai/Vocal-burst
We captioned them using Gemini Flash Audio 2.0. This dataset contains, this dataset contains ~ 365,000 vocal bursts from all kinds of categories.
It might be helpful for pre-training audio text foundation models to generate and understand all kinds of nuances in vocal bursts.
vocal-burst-classification-v2
Vocal Burst Classification V2 — laion/vocal-burst-classification-v2
The V2 training corpus for the Vocal Burst Classifier V2:
a single-label dataset over an 83-class vocal-burst taxonomy (82 non-speech human vocalizations
no_burst, index 82), shipped as precomputed VoiceCLAP-commercial embeddings plus the raw
vocal-bursts-clean audio.
The vocal-burst clips were generated with various synthetic text-to-audio models such as DramaBox,
then annotated and filtered as described… See the full description on the dataset page: https://huggingface.co/datasets/laion/vocal-burst-classification-v2.vocal-bursts
Vocal Bursts
A curated collection of 28,564 non-speech vocal burst audio samples across 18 categories.
Categories
Category
Samples
Breath
1,690
Cough
4,248
Crying
2,020
Laughter
4,797
Lip Popping
54
Lip Smacking
45
Moan
45
Nose Blowing
48
Pant
44
Scream
714
Sigh
3,546
Sneeze
3,813
Sniff
3,504
Teeth Chattering
46
Teeth Grinding
43
Throat Clearing
3,549
Tongue Clicking
47
Yawn
311
Format
Audio: FLAC
Metadata:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/vocal-bursts.vocal-bursts-taxonomy-DACVAE
Vocal Bursts Taxonomy — DACVAE + MaestroClap Embeddings & Scores
Processed version of with DACVAE latents, MaestroClap embeddings, derived attribute/quality/speaker scores, and Gemini-verified labels.
Overview
Metric
Value
Total samples
16,175
Categories
82
Genders
male, female
Female samples
8,097
Male samples
8,078
Gemini Label Verification
Every sample was sent to Gemini 3.1 Flash Lite for two independent tasks:
Match scoring:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/vocal-bursts-taxonomy-DACVAE.vocal_bursts_taxonomy_100_clean_wdsvocal-burst-annotation-asr-tuning-dataset
Vocal Burst Annotation ASR Tuning Dataset
A synthetic 500,000-sample multilingual dataset for training ASR models with inline vocal burst captioning, speaker diarization, and sentence-level timestamps. Each sample is approximately 1 minute of audio containing speech segments interleaved with vocal bursts (laughs, sighs, coughs, etc.), annotated with precise timing information.
Example Transcript
[nasalized, affirmative hum, steady pitch, moderate intensity]… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/vocal-burst-annotation-asr-tuning-dataset.more-synthetic-vocalbursts-raw
More Synthetic Vocal Bursts (Raw)
Synthetic vocal burst audio samples generated from a taxonomy of 202 vocal burst types across multiple text-to-audio and TTS models. Each sample is a short (3–10 second) non-speech vocalization — laughs, cries, gasps, sighs, growls, etc. — generated from text prompts describing the burst type, gender, and age group.
Models Used
Model
Type
Samples
Sample Rate
Notes
DramaBox (ResembleAI/Dramabox)
TTS DiT
2000
44.1 kHz… See the full description on the dataset page: https://huggingface.co/datasets/laion/more-synthetic-vocalbursts-raw.vocal-bursts
Vocal Bursts
A curated collection of 28,564 non-speech vocal burst audio samples across 18 categories.
Categories
Category
Samples
Breath
1,690
Cough
4,248
Crying
2,020
Laughter
4,797
Lip Popping
54
Lip Smacking
45
Moan
45
Nose Blowing
48
Pant
44
Scream
714
Sigh
3,546
Sneeze
3,813
Sniff
3,504
Teeth Chattering
46
Teeth Grinding
43
Throat Clearing
3,549
Tongue Clicking
47
Yawn
311
Format
Audio: FLAC… See the full description on the dataset page: https://huggingface.co/datasets/0x3/vocal-bursts.vocal-bursts-clean-reannotated
vocal-bursts-clean — blind Gemini re-annotation subset
A re-annotated subset of laion/vocal-bursts-clean. Up to 500 samples per original class (68 classes) were re-labelled blind by Gemini 2.5 Flash Lite (non-thinking): the model received only the audio plus the full 82-class VocalBurst taxonomy and chose one category — it was never shown the original label.
Samples: 18348 · Original classes: 68 · Taxonomy: 82 classes (LAION VocalBurst)
Agreement with original label:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/vocal-bursts-clean-reannotated.vocal_bursts_taxonomy_100wds_vocal_burst_100synthetic-vocal-burstsgemini_vocal_bursts_speedPrompts rephrased with Qwen3 32B on Cerebras through HF Inference Providers
This version is the same as the previous version but with speed prompts.
vocal_bursts_taxonomyvocal_bursts_largeChristoph from LAION's annotations on LAION's Got Talent dataset, including only verified vocal bursts
gemini_vocal_burstsvocal_bursts_taxonomy_100_cleanopenai_vocal_burstsvocal_burstsopenai_vocal_bursts_rewordedPrompts rephrased with Qwen3 Next 80B A3B on Parasail
