datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vctk
VCTK
This is a processed clone of the VCTK dataset with leading and trailing silence removed using Silero VAD. A fixed 25 ms of padding has been added to both ends of each audio clip to (hopefully) imrprove training and finetuning.
The original dataset is available at: https://datashare.ed.ac.uk/handle/10283/3443.
Reproducing
This repository notably lacks a requirements.txt file. There's likely a missing dependency or two, but roughly:
pydub
tqdm
torch
torchaudio… See the full description on the dataset page: https://huggingface.co/datasets/jspaulsen/vctk.esd
Emotional Speech Dataset (ESD)
The Emotional Speech Dataset (ESD) is a multilingual emotional speech corpus containing parallel recordings in English and Chinese across 5 emotions.
Dataset Details
Total samples: 35,000
Speakers: 20 (10 Chinese, 10 English)
Emotions: anger, happiness, neutral, sadness, surprise (7,000 each)
Languages: Chinese (zh), English (en) - 17,500 each
Gender: 10 male, 10 female speakers
Dataset Structure
Column
Description… See the full description on the dataset page: https://huggingface.co/datasets/jspaulsen/esd.LatinAccents
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/jstack32/LatinAccents.youtube_watch_XDw2cw6Fh8g
