datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yodas2_sidon
YODAS2-Sidon
Overview
This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling.
YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks.
We resampled original sidon output to 24kHz due to a storage constraints.
The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.Emilia-YODAS-ENemilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset.
https://huggingface.co/datasets/amphion/Emilia-Dataset
behavior-sd
🎙️ Behavior-SD
Official repository for our NAACL 2025 paper:Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language ModelsSehun Lee*, Kang-wook Kim*, Gunhee Kim (* Equal contribution)
🏆 SAC Award Winner in Speech Processing and Spoken Language Understanding
🔗 Links
Project Page
Code
📖 Overview
We explores how to generate natural, behaviorally-rich full-duplex spoken dialogues using large language models (LLMs).
We introduce:… See the full description on the dataset page: https://huggingface.co/datasets/yhytoto12/behavior-sd.YouTube-Cantonese
Cantonese Audio Dataset from YouTube
This dataset contains Cantonese audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models.
Data Source and Processing
The data was obtained through the following process:
Download: Audio (.m4a) and available Cantonese subtitles (.srt for zh-TW, zh-HK, zh-Hant)… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-Cantonese.Emilia-YODAS-DEM6common_voice_21_0_yuecantonese only
TMD2-DataYouTube-English
English Audio Dataset from YouTube
This dataset contains English audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models.
Data Source and Processing
The data was obtained through the following process:
Download: Audio (.m4a) and available English subtitles (.srt for en, en.j3PyPqV-e1s) were… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-English.Emilia-YODAS-KO-filteredemilia-yodas-alignedCLAP_freesound
LAION-Audio-630K Freesound Dataset
LAION-Audio-630K is the largest audio-text dataset publicly available and a magnitude larger than previous audio-text datasets (by 2022-11-05). Notably, it combines eight distinct datasets, which includes the Freesound dataset.
Specifically, this Hugging face repository contains two versions of Freesound dataset. Details of each dataset (e.g. how captions are made etc.) could be found in the "datacard" column of the table below.
Freesound (full):… See the full description on the dataset page: https://huggingface.co/datasets/YuXuAN0622/CLAP_freesound.jan2audiobEmilia-YODAS-KOEmilia-YODAS-FRYoutube-4Myoutube_ka_rajesh_raw_tempnsfw_tts_datasetA high-quality audio dataset designed for training and fine-tuning NSFW TTS models, including 30 characters, over 1000 hours of audio, and rich emotion/sound annotations.
Sample format: WAV (audio) + TXT (annotations), including emotion_label, sound_label and text.
Annotations: 6000+ emotion labels (intimate, breathy, teasing, etc.) and 760+ sound labels (moan, sigh, laugh, etc.) in the full version.
Audio sample is as follows:
[intimate, breathy, pleased] Oh, <moan> it feels so good when your… See the full description on the dataset page: https://huggingface.co/datasets/DMC-ykfx33/nsfw_tts_dataset.Emilia-YODAS-JAyoutube_kn_rajesh_raw_tempscam-nonscam-youtube-callsyoutube_ka_bookbrahma_raw_tempCommon_Voice_Delta_Segment_11.0youtube_te_dataset_raw_tempyoutube_ka_totalkannadamedia_raw_tempasc_testset
Yougen/asc_testset
Audio Scene Classification (ASC) speech dataset, packed as WebDataset tar shards.
Layout
data/
train/
metadata.csv
audio/
train-000.tar
train-001.tar
...
validation/
metadata.csv
audio/
validation-000.tar
...
test/
metadata.csv
audio/
test-000.tar
...
Shard counts:
test_a1: 8 tar shard(s)
test_a2: 16 tar shard(s)
test_a3: 13 tar shard(s)
test_a4: 15 tar shard(s)
test_a5: 8 tar… See the full description on the dataset page: https://huggingface.co/datasets/Yougen/asc_testset.Emilia-YODAS-ZHpfann_dataset
Yougen/pfann_dataset
Pfann-style audio dataset, packed as WebDataset tar shards.
Layout
data/
audioset/
metadata.csv
audio/
audioset-000.tar
...
fma_large/
metadata.csv
audio/
fma_large-000.tar
...
micirp/
metadata.csv
audio/
micirp-000.tar
...
Shard counts:
audioset: 3 tar shard(s)
fma_large: 136 tar shard(s)
micirp: 1 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:… See the full description on the dataset page: https://huggingface.co/datasets/Yougen/pfann_dataset.
