genuineness
voice-pairs-genuineness-vocalburst-blend-human-eval
Voice pairs for human evaluation of the genuineness and vocal-burst-blend predictors
This dataset exists to be annotated by humans. It is not a training set. Its only purpose is
to let human listeners adjudicate two automatic predictors by presenting them with the comparisons
those predictors claim to be able to make:
predictor
what it scores
native scale
repo
genuineness
how much a clip sounds like a real, lived-in spoken moment rather than a rehearsed or synthetic… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-pairs-genuineness-vocalburst-blend-human-eval.speech-genuineness
Speech Genuineness
~9,695 speech clips sampled from a variety of TTS-generated speech datasets and from the
Emolia corpus, each rated 0-6 for GENUINE NATURALNESS (how much it sounds like a real,
spontaneous, lived-in moment vs. a rehearsed/synthetic read) by Gemini-3.1-Pro.
Each clip is a 24 kHz mono MP3 of at most ~20 seconds. Clips carry an opaque id and an
anonymized bucket group id; the underlying source datasets are intentionally not identified.
The GENU… See the full description on the dataset page: https://huggingface.co/datasets/laion/speech-genuineness.
