datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenDialog
OpenDialog
OpenDialog is a 6.8k hours spoken dialogue dataset, introduced in the paper ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching.
Paper: https://arxiv.org/abs/2507.09318
GitHub: https://github.com/k2-fsa/ZipVoice
Project Page: https://zipvoice-dialog.github.io
OpenDialog is the first large-scale (6.8k hours) open-source spoken dialogue dataset derived from in-the-wild speech data. It consists of:
English data: 5074 hours
Chinese data: 1759… See the full description on the dataset page: https://huggingface.co/datasets/k2-fsa/OpenDialog.OpenS2S_Datasets
How to Use?
Download, merge the files, and extract
You can run the following command to merge the compressed file parts after downloading.
cat en_response_wav.tar.gz.* > en_response_wav.tar.gz
cat zh_response_wav.tar.gz.* > zh_response_wav.tar.gz
mid-classical-openmusenet4-3mbccfqwhoiswho_opensilero_open_stt
Usage
from datasets import load_dataset, Audio
dataset = load_dataset("Sh1man/silero_open_stt", "asr_calls_v2", split="train")
print(dataset[0]['wav'])
Subsets
The dataset contains three subsets:
asr_calls_v2: calls recordings
buriy_audio_books_2: books recordings
public_youtube700: youtube recordings
📊 Сводная статистика аудио-датасетов
📈 Общая статистика по всем датасетам
Метрика
Значение
Всего датасетов/сабсетов
3
Всего семплов… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/silero_open_stt.silero_open_stt_opus
Description
only subset tts_russian_addresses_rhvoice_4voices
Usage
from datasets import load_dataset, Audio
dataset = load_dataset("Sh1man/silero_open_stt_opus", "tts_russian_addresses_rhvoice_4voices", split="train")
print(dataset[0]['opus'])
Subsets
The dataset contains three subsets:
tts_russian_addresses_rhvoice_4voices: address recordings
📊 Сводная статистика аудио-датасетов
Информация по сплитам
🔹… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/silero_open_stt_opus.CapTTS-SFT-Audio
CapTTS-SFT Audio
DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details.
Overview
🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and text-to-speech… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapTTS-SFT-Audio.openslr_enhancedCapSpeech-MLS
CapSpeech-MLS Audio
DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details.
Overview
🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and text-to-speech… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech-MLS.CapSpeech-PT-SEDB-Audio
CapSpeech-PT-SEDB Audio
DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details.
Overview
🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech-PT-SEDB-Audio.openvid_ttsCapSpeech_Emilia
CapSpeech-Emilia Audio
DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details.
Overview
🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech_Emilia.CapSpeech-CommonVoice
CapSpeech-CommonVoice Audio
DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details.
Overview
🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech-CommonVoice.CapSpeech-PT-SEDB-HQ-Audio
CapSpeech-PT-SEDB-HQ Audio
DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details.
Overview
🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech-PT-SEDB-HQ-Audio.Emilia-ZHCapSpeech_GigaSpeech
CapSpeech-GigaSpeech Audio
DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details.
Overview
🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech_GigaSpeech.
