deepvk/NonverbalTTS
NonverbalTTS Dataset ๐ต๐ฃ๏ธ NonverbalTTS is a 17-hour open-access English speech corpus with aligned text annotations for nonverbal vocalizations (NVs) and emotional categories, designed to advance expressive text-to-speech (TTS) research. Key Features โจ 17 hours of high-quality speech data 10 NV types: Breathing, laughter, sighing, sneezing, coughing, throat clearing, groaning, grunting, snoring, sniffing 8 emotion categories: Angry, disgusted, fearful, happyโฆ See the full description on the dataset page: https://huggingface.co/datasets/deepvk/NonverbalTTS.
NonverbalTTS Dataset ๐ต๐ฃ๏ธ
  
NonverbalTTS is a 17-hour open-access English speech corpus with aligned text annotations for nonverbal vocalizations (NVs) and emotional categories, designed to advance expressive text-to-speech (TTS) research.
Key Features โจ
- 17 hours of high-quality speech data
- 10 NV types: Breathing, laughter, sighing, sneezing, coughing, throat clearing, groaning, grunting, snoring, sniffing
- 8 emotion categories: Angry, disgusted, fearful, happy, neutral, sad, surprised, other
- Diverse speakers: 2296 speakers (60% male, 40% female)
- Multi-source: Derived from VoxCeleb and Expresso corpora
- Rich metadata: Emotion labels, NV annotations, speaker IDs, audio quality metrics
- Sampling rate: 16kHz for audio from VoxCeleb, 48kHz for audio from Expresso <!-- ## Dataset Structure ๐
NonverbalTTS/ โโโ wavs/ # Audio files (16-48kHz WAV format) โ โโโ ex01sad00265.wav โ โโโ ... โโโ .gitattributes โโโ README.md โโโ metadata.csv # Metadata annotations -->
<!-- ## Metadata Schema (metadata.csv) ๐
<!-- NV Symbols: ๐ฌ๏ธ=Breath, ๐=Laughter, etc. (See Annotation Guidelines) -->
Loading the Dataset ๐ป
from datasets import load_dataset
dataset = load_dataset("deepvk/NonverbalTTS")<!-- # Access train split
# Output: {'index': 'ex01_sad_00265', 'file_name': 'wavs/ex01_sad_00265.wav', ...}
-->
## Annotation Pipeline ๐ง
1. **Automatic Detection**
- NV detection using [BEATs](https://arxiv.org/abs/2409.09546)
- Emotion classification with [emotion2vec+](https://huggingface.co/emotion2vec/emotion2vec_plus_large)
- ASR transcription via Canary model
2. **Human Validation**
- 3 annotators per sample
- Filtered non-English/multi-speaker clips
- NV/emotion validation and refinement
3. **Fusion Algorithm**
- Majority voting for final transcriptions
- Pyalign-based sequence alignment
- Multi-annotator hypothesis merging
## Benchmark Results ๐
Fine-tuning CosyVoice-300M on NonverbalTTS achieves parity with state-of-the-art proprietary systems:
|Metric | NVTTS | CosyVoice2 |
| ------- | ------- | ------- |
|Speaker Similarity | 0.89 | 0.85 |
|NV Jaccard | 0.8 | 0.78 |
|Human Preference | 33.4% | 35.4% |
## Use Cases ๐ก
- Training expressive TTS models
- Zero-shot NV synthesis
- Emotion-aware speech generation
- Prosody modeling research
## License ๐
- Annotations: CC BY-NC-SA 4.0
- Audio: Adheres to original source licenses (VoxCeleb, Expresso)
## Citation ๐
@inproceedings{borisov25_ssw, title = {{NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech}}, author = {Maksim Borisov and Egor Spirin and Daria Diatlova}, year = {2025}, booktitle = {{13th edition of the Speech Synthesis Workshop}}, pages = {104--109}, doi = {10.21437/SSW.2025-16}, }
