datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vctk
VCTK
This is a processed clone of the VCTK dataset with leading and trailing silence removed using Silero VAD. A fixed 25 ms of padding has been added to both ends of each audio clip to (hopefully) imrprove training and finetuning.
The original dataset is available at: https://datashare.ed.ac.uk/handle/10283/3443.
Reproducing
This repository notably lacks a requirements.txt file. There's likely a missing dependency or two, but roughly:
pydub
tqdm
torch
torchaudio… See the full description on the dataset page: https://huggingface.co/datasets/jspaulsen/vctk.esd
Emotional Speech Dataset (ESD)
The Emotional Speech Dataset (ESD) is a multilingual emotional speech corpus containing parallel recordings in English and Chinese across 5 emotions.
Dataset Details
Total samples: 35,000
Speakers: 20 (10 Chinese, 10 English)
Emotions: anger, happiness, neutral, sadness, surprise (7,000 each)
Languages: Chinese (zh), English (en) - 17,500 each
Gender: 10 male, 10 female speakers
Dataset Structure
Column
Description… See the full description on the dataset page: https://huggingface.co/datasets/jspaulsen/esd.lekol-lakay
Multi-Domain Haitian_creole Speech Dataset
This dataset contains 178 audio recordings with corresponding text transcriptions across multiple languages and domains.
Dataset Description
A comprehensive collection of audio files paired with text transcriptions, featuring both synthetic and natural speech across various domains. Suitable for automatic speech recognition (ASR), text-to-speech (TTS), and domain-specific speech processing tasks.
Dataset Structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/jsbeaudry/lekol-lakay.youtube_watch_XDw2cw6Fh8g
