CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Penaldo-CR7 /PenaldoCR7audion<1K0 likes4.3k downloads1mo agoHugging Face02cheelam /pendakwah_teknologi_yt_stt_datasetaudio100K<n<1M0 likes441 downloads2y agoHugging Face03penguinfish1688 /DuDE-Stage-III DuDE-Stage-III Training data for synchronized duplex speech modeling, derived from Seamless Interaction by Meta. full_conversations contains 7,387 complete conversations (488.973 conversation-hours; 977.946 participant-track hours), with 7,190 train and 197 dev examples. Both participant tracks retain their original shared clock. The longest example is 2,008 seconds. Internal ASR/codec chunks were reassembled; there is no example duration cap. short_windows preserves the earlier… See the full description on the dataset page: https://huggingface.co/datasets/penguinfish1688/DuDE-Stage-III.audiotext-to-speech10K<n<100K0 likes421 downloads1d agoHugging Face04skilledu /pendakwah_teknologi_yt_stt_datasetaudio100K<n<1M0 likes328 downloads3mo agoHugging Face05pengyizhou /nsc-imda-part6audio100K<n<1M0 likes286 downloads3mo agoHugging Face06pengyizhou /gigaspeech_subset_270haudio100K<n<1M0 likes270 downloads8mo agoHugging Face07pengine /Libritts_p_dataset_20260129 Contribution This dataset is a processed version of the original LibriTTS-P dataset, optimized for use on the Hugging Face platform. I've uploaded this version to make it more accessible to the community. All credit for the original data goes to the creators of LibriTTS-P. Changes Make a new column combined_prompt. The combined_prompt is a concatenation of the style_prompt and speaker_prompt, using the connector: "The speaker's identity can be described as ". In the… See the full description on the dataset page: https://huggingface.co/datasets/pengine/Libritts_p_dataset_20260129.audio100K<n<1M1 likes259 downloads8mo agoHugging Face08pengyizhou /accented_englishaudio100K<n<1M1 likes229 downloads3mo agoHugging Face09pengine /LIbritts_p_dataset_20260127 Contribution This dataset is a processed version of the original LibriTTS-P dataset, optimized for use on the Hugging Face platform. I've uploaded this version to make it more accessible to the community. All credit for the original data goes to the creators of LibriTTS-P. Changes Make a new column combined_prompt. The combined_prompt is a concatenation of the style_prompt and speaker_prompt, using the connector: "The speaker's identity can be described as ". In the… See the full description on the dataset page: https://huggingface.co/datasets/pengine/LIbritts_p_dataset_20260127.audio100K<n<1M0 likes211 downloads8mo agoHugging Face10pengyizhou /wenetspeech-subset-Saudio100K<n<1M0 likes175 downloads1y agoHugging Face11peng7554 /DS3500 English 中文 I. Basic Information of the Dataset Dataset Name: Underwater Acoustic Target Radiated Noise Dataset (including the original ShipsEar dataset and the enhanced DS3500 dataset) Dataset Version: V1.0 Release Date: July 2025 (based on the paper submission date) Update Records: First release, no updates yet Source and Contributors: Original ShipsEar dataset: Collected along the Atlantic coast of Spain from 2012 to 2013 Enhanced DS3500 dataset: Generated by institutions such… See the full description on the dataset page: https://huggingface.co/datasets/peng7554/DS3500.audio1K<n<10K2 likes80 downloads1y agoHugging Face12soiz1 /penguinmod-vm-prain PenguinMod/PenguinMod-Vm Modified Scratch VM with a JIT compiler and more features. This is a drop-in replacement for LLK/scratch-vm. Setup See https://github.com/TurboWarp/scratch-gui/wiki/Getting-Started to setup the complete TurboWarp environment. If you just want to play with the VM then it's the same process as upstream scratch-vm. Extension authors If you only use the standard reporter, boolean, and command block types, everything should just work without… See the full description on the dataset page: https://huggingface.co/datasets/soiz1/penguinmod-vm-prain.audion<1K0 likes67 downloads1y agoHugging Face13pengyichen /NaiLong-Voice-Clone 奶龙语音克隆数据集 完整项目与 Demo 效果可参见 GitHub 如果这个数据集对你有帮助,欢迎在 GitHub 上点个 Star ⭐ 支持一下! 数据集介绍 数据集按处理阶段分为以下四部分: 1. raw_audio (原始采样) 处理方式:使用 Audacity 直接对视频素材进行录音,格式为 44.1kHz, 16-bit, Stereo。 说明:包含背景音、特效及多角色对话的非结构化原片素材,是整个流水线的起点。 2. vocal_only (人声分离) 处理方式:从 raw_audio 中使用 UVR5 的 MDX-Net 模型剥离背景音乐与噪音。 说明:利用 MDX-Net 模型提取出干净的人声轨道,为后续切片提供高信噪比素材。 3. sliced_vocal (自动化切片) 处理方式:基于停顿检测、音色突变及总时长控制,将 vocal_only 自动化切分为一系列短音频。… See the full description on the dataset page: https://huggingface.co/datasets/pengyichen/NaiLong-Voice-Clone.audio10K<n<100K1 likes36 downloads5mo agoHugging Face14penguinfish1688 /DuDE-Stage-II DuDE-Stage-II Committed Stage II teacher self-distillation records: 337,389 train utterances (703.460 hours of synthetic codec targets) and 128 validation utterances. Each row is one independent utterance. The trainer samples independent utterances into two densely interleaved channels; this dataset does not assemble conversations. Targets were generated by Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice (weights revision 0c0e3051f131929182e2c023b9537f8b1c68adfe, upstream code revision… See the full description on the dataset page: https://huggingface.co/datasets/penguinfish1688/DuDE-Stage-II.tabulartext-to-speech100K<n<1M0 likes23 downloads1d agoHugging Face15pengyizhou /IALP-2026-data IALP-2026: Whisper Open-Set Data-Selection — Query / Dev / Test Sets Supporting data for the study "Whisper-Based Open-Set Data Selection for NSC Adaptation." This repository holds the fixed target-query, validation, and evaluation sets used across all experiments. Each part is a self-contained .tar.gz. All audio is 16 kHz mono. Each split ships with: audio/ — audio files (FLAC, except GigaSpeech which is WAV PCM_16) wav.scp — <utt_id> audio/<file> (Kaldi-style, relative paths)… See the full description on the dataset page: https://huggingface.co/datasets/pengyizhou/IALP-2026-data.audioautomatic-speech-recognition10K<n<100K0 likes20 downloads3mo agoHugging Face16pengyizhou /ESD-Subsetsaudio1K<n<10K0 likes19 downloads4mo agoHugging Face17pengyizhou /ESD-Subsets_answeraudio1K<n<10K0 likes15 downloads4mo agoHugging Face18Voice-man-76 /Pennyaudiotext-classificationn<1K0 likes7 downloads3y agoHugging Face19pengyizhou /NTU-refined-NSC-2021-evaluationaudio1K<n<10K0 likes6 downloads1y agoHugging Face20Stud1os /Daniel_Peninaudion<1K0 likes5 downloads3y agoHugging Face21penghor315 /custom-datasetaudion<1K0 likes5 downloads1y agoHugging Face22Shubhangi7 /tara_penaltygatedaudio10K<n<100K0 likes5 downloads4mo agoHugging Face23fdgvjhb /penny456audion<1K0 likes4 downloads3y agoHugging Face24pengyizhou /SWB1gatedaudio100K<n<1M0 likes4 downloads1y agoHugging Face25Shubhangi7 /simran_penaltygatedaudio10K<n<100K0 likes4 downloads4mo agoHugging Face26MoonIcee /peneraabritaaudion<1K0 likes2 downloads2y agoHugging Face27pengyizhou /hub5_english_eval_2000_swb1gatedaudioautomatic-speech-recognition1K<n<10K0 likes2 downloads1y agoHugging Face28pengyizhou /EN-MALAY-CSgatedaudio10K<n<100K0 likes2 downloads1y agoHugging Face29pengyizhou /EN-INDON-CSgatedaudio10K<n<100K0 likes1 downloads1y agoHugging Face30ardavey /emotion-penyisihan_satria_data_2025-datasetgatedaudion<1K1 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.