CoolFace
Datasetpublic

nytopop/expresso-conversational

The Expresso Dataset [paper] [demo samples] [Original repository] Introduction The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised). The transcriptions of the read speech are also provided. You can… See the full description on the dataset page: https://huggingface.co/datasets/nytopop/expresso-conversational.

sourceHugging Facecc-by-nc-4.0updated 1y agoView on Hugging Face
14likes418downloads
Dataset Card

The Expresso Dataset

[[paper]](https://arxiv.org/abs/2308.05725) [[demo samples]](https://speechbot.github.io/expresso/) [[Original repository]](https://github.com/facebookresearch/textlesslib/tree/main/examples/expresso/dataset)

Introduction

The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised). The transcriptions of the read speech are also provided.

You can listen to samples from the Expresso Dataset at this website.

Transcription & Segmentation

This subset of Expresso was machine transcribed and segmented with Parakeet TDT 0.6B V2.

Some transcription errors are to be expected.

Conversational Speech

The conversational config contains all improvised conversational speech as segmented 48KHz mono WAV.

Data Statistics

Here are the statistics of Expresso’s expressive styles:


StyleRead (min)Improvised (min)total (hrs)
angry-821.4
animal-270.4
animal_directed-320.5
awe-921.5
bored-921.5
calm-931.6
child-280.4
child_directed-380.6
confused94662.7
default1331584.9
desire-921.5
disgusted-1182.0
enunciated116623.0
fast-981.6
fearful-981.6
happy74922.8
laughing941033.3
narration21761.6
non_verbal-320.5
projected-941.6
sad811013.0
sarcastic-1061.8
singing*-4.07
sleepy-931.5
sympathetic-1001.7
whisper79862.8
Total11.5h34.4h45.9h

*singing is the only improvised style that is not in dialogue format.

Audio Quality

The audio was recorded in a professional recording studio with minimal background noise at 48kHz/24bit. The files for read speech and singing are in a mono wav format; and for the dialog section in stereo (one channel per actor), where the original flow of turn-taking is preserved.

License

The Expresso dataset is distributed under the CC BY-NC 4.0 license.

Reference

For more information, see the paper: EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis, Tu Anh Nguyen, Wei-Ning Hsu, Antony D'Avirro, Bowen Shi, Itai Gat, Maryam Fazel-Zarani, Tal Remez, Jade Copet, Gabriel Synnaeve, Michael Hassid, Felix Kreuk, Yossi Adi⁺, Emmanuel Dupoux⁺, INTERSPEECH 2023.