CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01k2-fsa /OpenDialog OpenDialog OpenDialog is a 6.8k hours spoken dialogue dataset, introduced in the paper ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching. Paper: https://arxiv.org/abs/2507.09318 GitHub: https://github.com/k2-fsa/ZipVoice Project Page: https://zipvoice-dialog.github.io OpenDialog is the first large-scale (6.8k hours) open-source spoken dialogue dataset derived from in-the-wild speech data. It consists of: English data: 5074 hours Chinese data: 1759… See the full description on the dataset page: https://huggingface.co/datasets/k2-fsa/OpenDialog.audiotext-to-speech100K<n<1M24 likes550 downloads5mo agoHugging Face02CASIA-LM /OpenS2S_Datasets How to Use? Download, merge the files, and extract You can run the following command to merge the compressed file parts after downloading. cat en_response_wav.tar.gz.* > en_response_wav.tar.gz cat zh_response_wav.tar.gz.* > zh_response_wav.tar.gz audio100K<n<1M8 likes461 downloads1y agoHugging Face03hidude562 /mid-classical-openmusenet4-3maudio10K<n<100K0 likes55 downloads9mo agoHugging Face0401gumano1d /bccfqwhoiswho_openaudio10K<n<100K0 likes31 downloads2y agoHugging Face05Sh1man /silero_open_stt Usage from datasets import load_dataset, Audio dataset = load_dataset("Sh1man/silero_open_stt", "asr_calls_v2", split="train") print(dataset[0]['wav']) Subsets The dataset contains three subsets: asr_calls_v2: calls recordings buriy_audio_books_2: books recordings public_youtube700: youtube recordings 📊 Сводная статистика аудио-датасетов 📈 Общая статистика по всем датасетам Метрика Значение Всего датасетов/сабсетов 3 Всего семплов… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/silero_open_stt.audio10K<n<100K4 likes29 downloads1y agoHugging Face06Sh1man /silero_open_stt_opus Description only subset tts_russian_addresses_rhvoice_4voices Usage from datasets import load_dataset, Audio dataset = load_dataset("Sh1man/silero_open_stt_opus", "tts_russian_addresses_rhvoice_4voices", split="train") print(dataset[0]['opus']) Subsets The dataset contains three subsets: tts_russian_addresses_rhvoice_4voices: address recordings 📊 Сводная статистика аудио-датасетов Информация по сплитам 🔹… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/silero_open_stt_opus.audio1M<n<10M0 likes22 downloads1y agoHugging Face07OpenSound /CapTTS-SFT-Audiogated CapTTS-SFT Audio DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details. Overview 🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and text-to-speech… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapTTS-SFT-Audio.audio100K<n<1M0 likes22 downloads1y agoHugging Face08Wikidepia /openslr_enhancedaudio100K<n<1M0 likes10 downloads7mo agoHugging Face09OpenSound /CapSpeech-MLSgated CapSpeech-MLS Audio DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details. Overview 🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and text-to-speech… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech-MLS.audio1M<n<10M1 likes8 downloads1y agoHugging Face10OpenSound /CapSpeech-PT-SEDB-Audiogated CapSpeech-PT-SEDB Audio DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details. Overview 🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech-PT-SEDB-Audio.audio1M<n<10M0 likes6 downloads1y agoHugging Face11Jordano /openvid_ttsaudio10K<n<100K0 likes6 downloads1y agoHugging Face12OpenSound /CapSpeech_Emiliagated CapSpeech-Emilia Audio DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details. Overview 🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech_Emilia.audio100K<n<1M3 likes5 downloads1y agoHugging Face13OpenSound /CapSpeech-CommonVoicegated CapSpeech-CommonVoice Audio DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details. Overview 🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech-CommonVoice.audio10K<n<100K1 likes4 downloads1y agoHugging Face14OpenSound /CapSpeech-PT-SEDB-HQ-Audiogated CapSpeech-PT-SEDB-HQ Audio DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details. Overview 🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech-PT-SEDB-HQ-Audio.audio100K<n<1M0 likes4 downloads1y agoHugging Face15OpenBot /Emilia-ZHaudio1M<n<10M0 likes4 downloads11mo agoHugging Face16OpenSound /CapSpeech_GigaSpeechgated CapSpeech-GigaSpeech Audio DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details. Overview 🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech_GigaSpeech.audio1M<n<10M0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.