CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amphion /Emilia-Datasetgated Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.audiotext-to-speech10M<n<100M489 likes46k downloads2y agoHugging Face02XRXRX /X-Voice-Dataset-Train X-Voice Training Dataset Overview The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling. Also the train set of X-Voice Model. Core Statistics Total Speech Duration: 420K hours 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.audiotext-to-speech10M<n<100M11 likes4.7k downloads5mo agoHugging Face03bhyuan /gptsovits_datasetgated bhyuan/gptsovits_dataset GPT-SoVITS speech dataset, packed as WebDataset tar shards. Layout data/ train/ metadata.csv audio/ train-000.tar train-001.tar ... validation/ metadata.csv audio/ validation-000.tar ... test/ metadata.csv audio/ test-000.tar ... Shard counts: youshengshu_v5_test: 6536 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/bhyuan/gptsovits_dataset.audiotext-to-speech10M<n<100M1 likes1.6k downloads4mo agoHugging Face04CASIA-LM /OpenS2S_Datasets How to Use? Download, merge the files, and extract You can run the following command to merge the compressed file parts after downloading. cat en_response_wav.tar.gz.* > en_response_wav.tar.gz cat zh_response_wav.tar.gz.* > zh_response_wav.tar.gz audio100K<n<1M8 likes461 downloads1y agoHugging Face05adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-hokkien Taiwan-Tongues-ASR-CE-dataset-hokkien 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hokkien.audioautomatic-speech-recognition10K<n<100K7 likes422 downloads9mo agoHugging Face06adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-zhtw Taiwan-Tongues-ASR-CE-dataset-zhtw 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。 🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-zhtw.audioautomatic-speech-recognition100K<n<1M2 likes208 downloads9mo agoHugging Face07farsi-asr /ganjoor-chunked-asr-datasetaudio100K<n<1M2 likes153 downloads2y agoHugging Face08itzune /antton-dataset Antton Dataset (Synthetic) This is a large-scale synthetic speech corpus designed for training and fine-tuning Basque Text-to-Speech (TTS) models. It consists of 99,996 audio files synthesized from the "Antton" voice model. This dataset was generated by Itzune and serves as the primary source for training the itzune/antton-tts (Piper version) model. Dataset Structure Due to the large volume of data (approx. 100,000 files), the dataset is organized in the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/itzune/antton-dataset.audiotext-to-speech100K<n<1M0 likes149 downloads6mo agoHugging Face09adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-hakka Taiwan-Tongues-ASR-CE-dataset-hakka 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。 🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hakka.audioautomatic-speech-recognition1K<n<10K2 likes125 downloads9mo agoHugging Face10arch-raven /music-fingerprint-dataset Neural Audio Fingerprint Dataset (c) 2021 by Sungkyun Chang https://github.com/mimbres/neural-audio-fp This dataset includes all music sources, background noise and impulse-reponses (IR) samples that have been used in the work ["Neural Audio Fingerprint for High-specific Audio Retrieval based on Contrastive Learning"] (https://arxiv.org/abs/2010.11910). Format: 16-bit PCM Mono WAV, Sampling rate 8000 Hz Description: / fingerprint_dataset_icassp2021/… See the full description on the dataset page: https://huggingface.co/datasets/arch-raven/music-fingerprint-dataset.audio10K<n<100K8 likes111 downloads4y agoHugging Face11issai /Multilingual_Speech_Dataset Multilingual Speech Dataset Paper: A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English Repository: https://github.com/IS2AI/MultilingualASR Description: This repository provides the dataset used in the paper "A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English". The paper focuses on training a single end-to-end (E2E) ASR model for Kazakh, Russian, and English, comparing monolingual and multilingual approaches… See the full description on the dataset page: https://huggingface.co/datasets/issai/Multilingual_Speech_Dataset.audioautomatic-speech-recognition100K<n<1M3 likes111 downloads2y agoHugging Face12farsi-asr /ganjoor-datasetaudio10K<n<100K0 likes93 downloads2y agoHugging Face13TTS-AGI /vocal-burst-annotation-asr-tuning-dataset Vocal Burst Annotation ASR Tuning Dataset A synthetic 500,000-sample multilingual dataset for training ASR models with inline vocal burst captioning, speaker diarization, and sentence-level timestamps. Each sample is approximately 1 minute of audio containing speech segments interleaved with vocal bursts (laughs, sighs, coughs, etc.), annotated with precise timing information. Example Transcript [nasalized, affirmative hum, steady pitch, moderate intensity]… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/vocal-burst-annotation-asr-tuning-dataset.audioautomatic-speech-recognition100K<n<1M2 likes92 downloads6mo agoHugging Face14cubbk /audio_swedish_2_dataset_cleanedaudio1K<n<10K0 likes56 downloads1y agoHugging Face15adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-en Taiwan-Tongues-ASR-CE-dataset-en 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。 🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-en.audioautomatic-speech-recognition10K<n<100K0 likes53 downloads9mo agoHugging Face16tts-dataset /japanese-singing-voicegated Japanese Singing Voice Dataset / 日本語歌声データセット English | 日本語 English A large-scale Japanese singing voice dataset for training voice conversion models. Dataset Description This dataset contains Japanese singing voice audio files collected for training singing voice conversion (SVC) models such as Seed-VC, RVC, So-VITS-SVC, and similar architectures. Dataset Statistics Metric Value Total Duration ~1,000 hours Number of Files 15,311… See the full description on the dataset page: https://huggingface.co/datasets/tts-dataset/japanese-singing-voice.audioaudio-classification10K<n<100K1 likes35 downloads9mo agoHugging Face17itzune /maider-dataset Maider Dataset (Synthetic) This is a large-scale synthetic speech corpus designed for training and fine-tuning Basque Text-to-Speech (TTS) models. It consists of 99,996 audio files synthesized from the "Maider" voice model. This dataset was generated by Itzune and serves as the primary source for training the itzune/maider-tts (Piper version) model. Dataset Structure Due to the large volume of data (approx. 100,000 files), the dataset is organized in the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/itzune/maider-dataset.audiotext-to-speech10K<n<100K0 likes30 downloads7mo agoHugging Face18TTS-AGI /voice-emo-cloning-dataset Emotion-Cloning TTS Training Dataset Location /home/deployer/laion/echo-tts-training-main/emotion_eval/dataset_output/ Overview This dataset contains ~22,518 training triplets for fine-tuning a zero-shot voice+emotion cloning TTS model. Each sample provides everything needed to train a model that can clone both a speaker's voice identity AND their emotional delivery from separate reference audio clips. The data is stored as WebDataset .tar shards, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/voice-emo-cloning-dataset.audio10K<n<100K0 likes27 downloads6mo agoHugging Face19DMC-ykfx33 /nsfw_tts_datasetgatedA high-quality audio dataset designed for training and fine-tuning NSFW TTS models, including 30 characters, over 1000 hours of audio, and rich emotion/sound annotations. Sample format: WAV (audio) + TXT (annotations), including emotion_label, sound_label and text. Annotations: 6000+ emotion labels (intimate, breathy, teasing, etc.) and 760+ sound labels (moan, sigh, laugh, etc.) in the full version. Audio sample is as follows: [intimate, breathy, pleased] Oh, <moan> it feels so good when your… See the full description on the dataset page: https://huggingface.co/datasets/DMC-ykfx33/nsfw_tts_dataset.texttext-to-speech10K<n<100K10 likes20 downloads6mo agoHugging Face20leo-fixie /jr-datasetaudio100K<n<1M0 likes18 downloads1y agoHugging Face21aegean-ai /engine-anomaly-detection-dataset license: other license_name: ntt license_link: https://zenodo.org/records/3351307/files/LICENSE.pdf?download=1 audio1K<n<10K3 likes17 downloads3y agoHugging Face22dalietng /dataset_asv_x2audio100K<n<1M0 likes16 downloads8mo agoHugging Face23datasets-examples /doc-audio-11 [doc] audio dataset 11 This dataset contains two tar files that contain pairs of samples with one audio file and one JSON file. audion<1K0 likes12 downloads2y agoHugging Face24cmeraki /youtube_te_dataset_raw_tempaudio1K<n<10K0 likes12 downloads2y agoHugging Face25nev /aitch-datasetaudion<1K0 likes9 downloads4y agoHugging Face26pavi1561 /Divehi_text_speech_datasetaudion<1K0 likes8 downloads2y agoHugging Face27tts-dataset /filtered-gol-datasetgated Filtered GOL Dataset midralab/gol-dataset をTTS(Text-to-Speech)学習用にフィルタリングしたデータセットです。 データセット概要 項目 値 総再生時間 約1,880時間 サンプル数 約120万 話者数 380人 データサイズ 約280GB 形式 WebDataset (.tar) 音声形式 FLAC (44.1kHz, モノラル) フィルタリング条件 基本フィルタ テキスト長: 3文字以上 音声長: 1秒以上、60秒未満 話者フィルタ 話者あたり5時間以上の音声データを持つ話者のみ テキストフィルタ(除外対象) 非言語テキスト(句読点のみ、空白のみなど) 顔文字 (^_^), (T_T) など 笑い表現 (笑), 文末の www 絵文字 英数字のみのテキスト 同一文字4回以上の繰り返し データ構造… See the full description on the dataset page: https://huggingface.co/datasets/tts-dataset/filtered-gol-dataset.audiotext-to-speech1M<n<10M1 likes8 downloads8mo agoHugging Face28sambal /lava_datasetaudio100K<n<1M0 likes6 downloads2y agoHugging Face29WariHima /wav2phone-datasetjvnv-nonvデータセットの話者f2で学習したVoiceSpeechMakerのモデルで、jsutコーパスの書き下し分を推論し、推論時のフルコンテキストラベルとペアにしたデータセットです フルコンテキストラベルは次のような形式です(hfのプレビューは壊れています) xx^xx-sil+m=i/A:xx+xx+xx/B:xx-xx_xx/C:xx_xx+xx/D:02+xx_xx/E:xx_xx!xx_xx-xx/F:xx_xx#xx_xx@xx_xx|xx_xx/G:3_3%0_0_xx/H:xx_xx/I:xx-xx@xx+xx&xx-xx|xx+xx/J:3_23/K:1+3-23 xx^sil-m+i=z/A:-2+1+3/B:xx-xx_xx/C:02_xx+xx/D:13+xx_xx/E:xx_xx!xx_xx-xx/F:3_3#0_0@1_3|1_23/G:7_2%0_0_1/H:xx_xx/I:3-23@1+1&1-3|1+23/J:xx_xx/K:1+3-23… See the full description on the dataset page: https://huggingface.co/datasets/WariHima/wav2phone-dataset.audio10K<n<100K2 likes6 downloads1y agoHugging Face30AlphJain /demo_tts_datasetaudion<1K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.