CoolFace
Datasetpublic

amphion/Debatts-Data

Debatts-Data: The First Madarin Rebuttal Speech Dataset for Expressive Text-to-Speech Synthesis The Debatts-Data dataset is the first Madarin rebuttal speech dataset for expressive text-to-speech synthesis. It is constructed from a vast collection of professional Madarin speech data sourced from diverse video platforms and podcasts on the Internet. The in-the-wild collection approach ensures the real and natural rebuttal speech. In addition, the dataset contains annotations of… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Debatts-Data.

sourceHugging Facecc-by-nc-4.0updated 2y agoView on Hugging Face
11likes71downloads
Dataset Card

Debatts-Data: The First Madarin Rebuttal Speech Dataset for Expressive Text-to-Speech Synthesis

The Debatts-Data dataset is the first Madarin rebuttal speech dataset for expressive text-to-speech synthesis. It is constructed from a vast collection of professional Madarin speech data sourced from diverse video platforms and podcasts on the Internet. The in-the-wild collection approach ensures the real and natural rebuttal speech. In addition, the dataset contains annotations of transcription, duration and style embed. The table and chart below provide the statistic information for the dataset. For some dataset samples and more information regarding Debatts system, please visit the Debatts project page.

Dataset Specifications

AttributeValue
LanguageZH
Number of Speakers2,350 (est.)
Duration (hrs)111
TypeText + Speech
Sample Rate (kHz)16
Recorded MethodWild

The JSON files in the dataset contain the following keys:

KeyDescription
keyUnique identifier for each sample in the dataset
textText transcription of the audio
durationDuration of the audio clip in seconds
languageLanguage of the audio content
wav_pathPath to the corresponding WAV file
prompt0_wav_pathPath to the WAV file used as a prompt
style_featureStyle features associated with the audio sample

README 🔥🔥🔥

Dataset Usage

To utilize the Debatts-Data dataset, you can download the raw audio files from the files and versions. The Debatts-Data.tar.gz contains the training data, while the Debatts-Data_test.tar.gz contains the testing data with extra speaker prompt speech.

Please note that Debatts-Data does not own the copyright to the audio files; the copyright remains with the original owners of the videos or audio. Users are permitted to use this dataset only for non-commercial purposes under the CC BY-NC-4.0 license.