datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.X-Voice-Dataset-Train
X-Voice Training Dataset
Overview
The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling.
Also the train set of X-Voice Model.
Core Statistics
Total Speech Duration: 420K hours
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.gptsovits_dataset
bhyuan/gptsovits_dataset
GPT-SoVITS speech dataset, packed as WebDataset tar shards.
Layout
data/
train/
metadata.csv
audio/
train-000.tar
train-001.tar
...
validation/
metadata.csv
audio/
validation-000.tar
...
test/
metadata.csv
audio/
test-000.tar
...
Shard counts:
youshengshu_v5_test: 6536 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/bhyuan/gptsovits_dataset.majestrino-datavoice-data
Voice-Data: a curated multi-corpus voice dataset for voice–text contrastive training
voice-data is a single, globally-shuffled WebDataset that bundles several voice/speech corpora into one ready-to-train mixture for voice–text contrastive (CLAP-style) models such as VoiceCLAP. Each clip pairs 48 kHz mono FLAC audio with a natural-language text caption describing the voice — its emotion, prosody, timbre, speaking style, recording context, and speaker traits.
The distinguishing… See the full description on the dataset page: https://huggingface.co/datasets/gijs/voice-data.Ola-DataThis repository contains the data presented in Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment.
Code: https://github.com/Ola-Omni/Ola
log_prompt_datavoiceclap-data
VoiceCLAP Data
The audio + dense-caption mixture used to train
laion/voiceclap-small and
laion/voiceclap-large.
Each tar shard is a WebDataset of
paired <key>.flac (48 kHz mono audio) + <key>.json (caption + metadata)
samples. Captions and structured attribute annotations are produced
automatically by a pipeline of audio-aware LLMs — Qwen-Audio, Gemini Flash 2.5,
and a thinking-mode reasoning model that scores emotion under the EmoNet
taxonomy plus per-clip vocal-burst, timbre… See the full description on the dataset page: https://huggingface.co/datasets/laion/voiceclap-data.OpenS2S_Datasets
How to Use?
Download, merge the files, and extract
You can run the following command to merge the compressed file parts after downloading.
cat en_response_wav.tar.gz.* > en_response_wav.tar.gz
cat zh_response_wav.tar.gz.* > zh_response_wav.tar.gz
german_mTaiwan-Tongues-ASR-CE-dataset-hokkien
Taiwan-Tongues-ASR-CE-dataset-hokkien
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hokkien.xares_llm_dataTaiwan-Tongues-ASR-CE-dataset-zhtw
Taiwan-Tongues-ASR-CE-dataset-zhtw
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。
🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-zhtw.ganjoor-chunked-asr-datasetantton-dataset
Antton Dataset (Synthetic)
This is a large-scale synthetic speech corpus designed for training and fine-tuning Basque Text-to-Speech (TTS) models. It consists of 99,996 audio files synthesized from the "Antton" voice model.
This dataset was generated by Itzune and serves as the primary source for training the itzune/antton-tts (Piper version) model.
Dataset Structure
Due to the large volume of data (approx. 100,000 files), the dataset is organized in the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/itzune/antton-dataset.TMD2-Dataspeech_dataTaiwan-Tongues-ASR-CE-dataset-hakka
Taiwan-Tongues-ASR-CE-dataset-hakka
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。
🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hakka.music-fingerprint-dataset
Neural Audio Fingerprint Dataset
(c) 2021 by Sungkyun Chang
https://github.com/mimbres/neural-audio-fp
This dataset includes all music sources, background noise and impulse-reponses
(IR) samples that have been used in the work ["Neural Audio Fingerprint for
High-specific Audio Retrieval based on Contrastive Learning"]
(https://arxiv.org/abs/2010.11910).
Format:
16-bit PCM Mono WAV, Sampling rate 8000 Hz
Description:
/
fingerprint_dataset_icassp2021/… See the full description on the dataset page: https://huggingface.co/datasets/arch-raven/music-fingerprint-dataset.Multilingual_Speech_Dataset
Multilingual Speech Dataset
Paper: A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English
Repository: https://github.com/IS2AI/MultilingualASR
Description: This repository provides the dataset used in the paper "A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English". The paper focuses on training a single end-to-end (E2E) ASR model for Kazakh, Russian, and English, comparing monolingual and multilingual approaches… See the full description on the dataset page: https://huggingface.co/datasets/issai/Multilingual_Speech_Dataset.ganjoor-datasetvocal-burst-annotation-asr-tuning-dataset
Vocal Burst Annotation ASR Tuning Dataset
A synthetic 500,000-sample multilingual dataset for training ASR models with inline vocal burst captioning, speaker diarization, and sentence-level timestamps. Each sample is approximately 1 minute of audio containing speech segments interleaved with vocal bursts (laughs, sighs, coughs, etc.), annotated with precise timing information.
Example Transcript
[nasalized, affirmative hum, steady pitch, moderate intensity]… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/vocal-burst-annotation-asr-tuning-dataset.voice-annotation-data-v2
Voice Annotation Data v2
A curated dataset of 18,632 audio samples (9,391 positives + 9,241 negatives) across 58 voice dimensions. Each bucket contains up to 25 positive examples (audio that clearly fits the bucket) and 25 negative examples (audio confirmed to NOT fit the bucket by Gemini 2.0 Flash).
Changes from v1
Positive + Negative pairs: Every bucket now has up to 25 confirmed negative examples alongside 25 positives
EXPL redefined: Content Appropriateness reduced… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/voice-annotation-data-v2.audio_swedish_2_dataset_cleanedTaiwan-Tongues-ASR-CE-dataset-en
Taiwan-Tongues-ASR-CE-dataset-en
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。
🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-en.uzbek_stt_dataspeech_datatsetjapanese-singing-voice
Japanese Singing Voice Dataset / 日本語歌声データセット
English | 日本語
English
A large-scale Japanese singing voice dataset for training voice conversion models.
Dataset Description
This dataset contains Japanese singing voice audio files collected for training singing voice conversion (SVC) models such as Seed-VC, RVC, So-VITS-SVC, and similar architectures.
Dataset Statistics
Metric
Value
Total Duration
~1,000 hours
Number of Files
15,311… See the full description on the dataset page: https://huggingface.co/datasets/tts-dataset/japanese-singing-voice.maider-dataset
Maider Dataset (Synthetic)
This is a large-scale synthetic speech corpus designed for training and fine-tuning Basque Text-to-Speech (TTS) models. It consists of 99,996 audio files synthesized from the "Maider" voice model.
This dataset was generated by Itzune and serves as the primary source for training the itzune/maider-tts (Piper version) model.
Dataset Structure
Due to the large volume of data (approx. 100,000 files), the dataset is organized in the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/itzune/maider-dataset.voice-emo-cloning-dataset
Emotion-Cloning TTS Training Dataset
Location
/home/deployer/laion/echo-tts-training-main/emotion_eval/dataset_output/
Overview
This dataset contains ~22,518 training triplets for fine-tuning a zero-shot voice+emotion cloning TTS model. Each sample provides everything needed to train a model that can clone both a speaker's voice identity AND their emotional delivery from separate reference audio clips.
The data is stored as WebDataset .tar shards, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/voice-emo-cloning-dataset.
