datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vocalvocal-affect-bench
VocalAffectBench
VocalAffectBench is a test-only benchmark for evaluating whether AI audio models can identify expressed vocal emotion from raw audio.
Paper: VocalAffectBench: Evaluating Vocal Emotion Recognition in AI Audio Models
The benchmark targets the expressed emotion — what the speaker conveys through vocal tone, prosody, pace, intensity, and pauses — not inferred internal state.
Contents
280 human-recorded English WAV clips, totalling 2.32 hours.
7… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/vocal-affect-bench.vocalcoachbench-review
VocalCoachBench
VocalCoachBench is a singing-audio benchmark for evaluating vocal coaching
judgments. This release contains expert annotations for 515 singing recordings:
free-form coaching feedback, atomic diagnosis/correction claims, Top-3 issue
labels, same-song triplet rankings, and segment-conditioned issue labels.
Subsets:
same_song / Dataset A: 207 Amazing Grace performances from DAMP-S-AG.
Audio is not redistributed; use audio_filename to match the official release.… See the full description on the dataset page: https://huggingface.co/datasets/vocalcoachbench/vocalcoachbench-review.speech2speech_vocalnethebrew-targum-vocalized
Vocalized Hebrew–Targum Parallel Corpus
Verse-aligned parallel corpus of Biblical Hebrew and Targumic Aramaic, covering 15119 verses of the Hebrew Bible.
Structure
field
type
description
book
int
Book number, 1–39 in standard Hebrew Bible order
book_name
string
English book name
chapter
int
Chapter
verse
int
Verse
hebrew
string
Masoretic Hebrew
targum
string
Targum Onkelos / Jonathan
split
verses
train
13646
validation
718… See the full description on the dataset page: https://huggingface.co/datasets/johnlockejrr/hebrew-targum-vocalized.Vocalset-Breath
Dataset Card — VocalSet-Breath
Dataset summary
VocalSet-Breath is a hand-labeled annotation layer on top of VocalSet (Wilkins et al., 2018) that adds time-aligned audible-breath-event labels. To our knowledge it is the first openly released dedicated breath-event layer for singing voice — onset/offset times with per-event confidence and explicit hard negatives. (Singing corpora such as GTSinger, Opencpop, and M4Singer carry breath only as phoneme tokens or… See the full description on the dataset page: https://huggingface.co/datasets/Ewakaa/Vocalset-Breath.Acoustic-Emotion-Vocal-Signature
Acoustic Predictors of Emotional States and Distress Signatures
Project Overview:
This dataset is sourced from Kaggle (Speech Emotion Detection Dataset), containing features from the RAVDESS and TESS databases.
The dataset consists of approximately 10,000 rows and 12 features.
Core Features: Key attributes include Pitch (Hz), Intensity (dB), and 13 coefficients of Mel-Frequency Cepstral Coefficients (MFCCs).
Clinical Application: This research facilitates the identification of a "vocal… See the full description on the dataset page: https://huggingface.co/datasets/yuvalhazan2/Acoustic-Emotion-Vocal-Signature.carnatic-raga-vocalsjapanese-singing-voice-vocal-only-audit
Japanese singing voice vocal-only — aggregate audit
This one-row audit describes tts-dataset/japanese-singing-voice-vocal-only at revision
c3ea48aa3909c23fae8e04ca28bed7ab84054066. It excludes titles, source names, URLs,
item IDs, paths, JSON metadata, hashes, and audio.
Sparse tar indexing found 4,871 complete JSON/WAV pairs across 40 unique archives, with
zero irregular pairs, unsafe paths, walk errors, or unterminated tars. Forty bounded WAV
samples were 44.1 kHz stereo… See the full description on the dataset page: https://huggingface.co/datasets/tts-dataset/japanese-singing-voice-vocal-only-audit.VocalVirtuoso
VocalVirtuoso
tags: generative, music, speech synthesis
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'VocalVirtuoso' dataset is a curated collection of high-fidelity audio samples featuring various professional vocalists, each selected for their ability to produce distinct musical genres and styles. The dataset is intended for research and development in the field of generative music and speech synthesis, particularly… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/VocalVirtuoso.
