CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sarulab-speech /yodas2_sidon YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.audiotext-to-speech1M<n<10M65 likes32k downloads10mo agoHugging Face02speechcolab /gigaspeech2gated Dataset Card for GigaSpeech 2 Dataset Description GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese. Repository: https://github.com/SpeechColab/GigaSpeech2 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.audioautomatic-speech-recognition10M<n<100M71 likes5.8k downloads6mo agoHugging Face03lighthouse-emnlp2024 /Clotho-Moment Clotho-Moment This repository provides wav files used in Language-based Audio Moment Retrieval. Each sample includes long audio containing some audio events with the temporal and textual annotation. Project page: https://h-munakata.github.io/Language-based-Audio-Moment-Retrieval/ Code: https://github.com/line/lighthouse Split Train train/train-{000..715}.tar 37930 audio samples Valid valid/valid-{000..108}.tar 5741 audio samples Test test/test-{000..142}.tar 7569… See the full description on the dataset page: https://huggingface.co/datasets/lighthouse-emnlp2024/Clotho-Moment.audioaudio-text-to-text10K<n<100K2 likes3.6k downloads8mo agoHugging Face04mitermix /audiosnippets_small_with_detailed_annotation2audio1M<n<10M1 likes2.7k downloads2y agoHugging Face05sarulab-speech /commonvoice22_sidongated CV22-Sidon Overview This dataset hosts a release of Mozilla Common Voice 22 restored with the Sidon speech restoration model. Source: Mozilla Common Voice 22.0 Processing: Sidon denoising (sarulab-speech/sidon-v0.1) with 21 s chunks and 48 kHz reconstruction Format: WebDataset shards (.tar.gz) Manifest: paths.yaml enumerates every shard path for Hugging Face–style loading License: Original Common Voice license (CC0 1.0) Languages 137 language folders are… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon.audiotext-to-speech10M<n<100M30 likes1.9k downloads1y agoHugging Face06krishnakalyan3 /emo_webds_2audio10K<n<100K7 likes1.6k downloads2y agoHugging Face07mitermix /audiosnippets_long_2_5Maudio1M<n<10M3 likes1.4k downloads2y agoHugging Face08TTS-AGI /commonvoice22-sidon-dacvae CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents Source sarulab-speech/commonvoice22_sidon Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.audioautomatic-speech-recognition1M<n<10M1 likes1.1k downloads6mo agoHugging Face09litagin /reazon-speech-v2-clonegated Reazon Speech v2 dataset mirror Original Dataset Source Hugging Face Dataset Page: reazon-research/reazonspeech Project Page: Reazon Research License This dataset is a mirror of the original Reazon Speech v2 dataset, but on 🤗 server (so may be faster). This dataset is licensed under the CDLA-Sharing-1.0. The original dataset comes with the following restriction: TO USE THIS DATASET, YOU MUST AGREE THAT YOU WILL USE THE DATASET SOLELY FOR THE PURPOSE OF… See the full description on the dataset page: https://huggingface.co/datasets/litagin/reazon-speech-v2-clone.audioautomatic-speech-recognition10K<n<100K12 likes822 downloads2y agoHugging Face10k2-fsa /OpenDialog OpenDialog OpenDialog is a 6.8k hours spoken dialogue dataset, introduced in the paper ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching. Paper: https://arxiv.org/abs/2507.09318 GitHub: https://github.com/k2-fsa/ZipVoice Project Page: https://zipvoice-dialog.github.io OpenDialog is the first large-scale (6.8k hours) open-source spoken dialogue dataset derived from in-the-wild speech data. It consists of: English data: 5074 hours Chinese data: 1759… See the full description on the dataset page: https://huggingface.co/datasets/k2-fsa/OpenDialog.audiotext-to-speech100K<n<1M24 likes550 downloads5mo agoHugging Face11sheng22213 /multi_round_speech_180kaudio1M<n<10M2 likes493 downloads1y agoHugging Face12ESpeech /ESpeech-webinars2 Webinar Audio Dataset Dataset Description This dataset contains 850 hours processed webinar audio segments with corresponding metadata. Each audio file represents a segment extracted from webinar recordings, processed at 44.1kHz sample rate. Dataset Summary Language: Russian Task: TTS, ASR, Quality Asessment Audio format: MP3, 44.1kHz sample rate Structure: Segmented audio files with JSON metadata Dataset Structure Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-webinars2.audiotext-to-speech100K<n<1M8 likes378 downloads1y agoHugging Face13laion /Emilia-with-Emotion-Annotations2audio10M<n<100M1 likes317 downloads1y agoHugging Face14humanify /ht2_44khzaudio1M<n<10M2 likes232 downloads6mo agoHugging Face15Sh1man /common_voice_21_ru Dataset Description Набор данных validated.tsv отфильтрованный по down_votes = 0 📊 Статистика датасета Информация по сплитам 🔹 Тренировочный набор (train) Метрика Значение Количество семплов 93,531 Общая продолжительность 132.25 часов (476,089.70 секунд) Средняя продолжительность семпла 5.09 секунд 🔹 Валидационный набор (validate) Метрика Значение Количество семплов 38,836 Общая продолжительность 55.21… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/common_voice_21_ru.audio100K<n<1M5 likes170 downloads1y agoHugging Face16sheng22213 /speech_text-tts_audioaudio10K<n<100K0 likes151 downloads1y agoHugging Face17acul3 /Audiobook_Noice_V2audio10K<n<100K0 likes130 downloads2y agoHugging Face18novateur /cosyvoice2_enaudio100K<n<1M0 likes105 downloads2y agoHugging Face19ming030890 /common_voice_21_0_yuecantonese only audio10K<n<100K0 likes101 downloads1y agoHugging Face20yangxiaoda /TMD2-Dataaudio100K<n<1M1 likes100 downloads4mo agoHugging Face21TTS-AGI /voice-annotation-data-v2 Voice Annotation Data v2 A curated dataset of 18,632 audio samples (9,391 positives + 9,241 negatives) across 58 voice dimensions. Each bucket contains up to 25 positive examples (audio that clearly fits the bucket) and 25 negative examples (audio confirmed to NOT fit the bucket by Gemini 2.0 Flash). Changes from v1 Positive + Negative pairs: Every bucket now has up to 25 confirmed negative examples alongside 25 positives EXPL redefined: Content Appropriateness reduced… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/voice-annotation-data-v2.audioaudio-classification10K<n<100K3 likes73 downloads5mo agoHugging Face22Darknsu /voxceleb2-40k-part1-preprocess-all-files-separateaudio100K<n<1M0 likes58 downloads5mo agoHugging Face23cubbk /audio_swedish_2_dataset_cleanedaudio1K<n<10K0 likes56 downloads1y agoHugging Face24AnodHuang /ASV_Spoof_2019_LA_SNR_50MBaudio100K<n<1M1 likes56 downloads9mo agoHugging Face25Viol2000 /musdb18hqaudion<1K0 likes55 downloads1y agoHugging Face26ai-music4you2 /ai-generated-songs2audio100K<n<1M0 likes48 downloads7mo agoHugging Face27backups /ai-m-2audio10M<n<100M2 likes44 downloads1y agoHugging Face28PuristanLabs1 /urdu-turn-detection-audio-v2 🗣️ Urdu Turn Detection (Audio Dataset V2) This is the official dataset for the model [PuristanLabs1/urdu-turn-v2](https://huggingface.co/PuristanLabs1/urdu-turn-v2), a high precision, low latency system for detecting the end of a conversational turn in Urdu speech. It contains 11,479 audio clips (balanced between Complete and Incomplete) specifically designed to train robust models for realtime Voice AI applications like "Smart Turn" or "Barge-in" detection. 🚀 How… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/urdu-turn-detection-audio-v2.audioaudio-classification10K<n<100K0 likes42 downloads9mo agoHugging Face29guangzhaoli /multilingual-test-distill-strong-tts-20260520 Multilingual Test Distill Strong TTS 20260520 This repository contains a distributable tar-sharded version of multilingual_test_distill_strong_tts_20260520. The dataset follows the local voice_dataset/data layout after extraction: data/csvs/metadata_zh.csv data/csvs/metadata_en.csv data/csvs/metadata_ja.csv data/csvs/metadata_ko.csv data/zh/**/*.wav data/en/**/*.wav data/ja/**/*.wav data/ko/**/*.wav Metadata format: file_path|duration|dnsmos|text dnsmos is intentionally blank… See the full description on the dataset page: https://huggingface.co/datasets/guangzhaoli/multilingual-test-distill-strong-tts-20260520.audiotext-to-speech1K<n<10K0 likes38 downloads4mo agoHugging Face30emoji-tts /emoji-tts-22k Emoji-TTS 22K Training Data Emoji-TTS 22K is the training corpus used to build Emoji-TTS, an emoji-conditioned expressive text-to-speech model. Each example pairs an English transcript, an emoji control label, and a synthetic WAV utterance spoken with the fixed Kore voice. The corpus contains 21,940 utterances: 19,945 emoji-conditioned samples and 1,995 neutral/no-emoji samples. The control inventory covers ten emoji labels plus the neutral <none> label. How It Was Built… See the full description on the dataset page: https://huggingface.co/datasets/emoji-tts/emoji-tts-22k.audiotext-to-speech10K<n<100K0 likes33 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.