CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mesolitica /Zeroshot-Audio-Classification-Instructions Zeroshot-Audio-Classification-Instructions Convert audio classification dataset into zero-shot format speech instructions, support both single label and multi-label, VGGSound FSD50k Nonspeech7k urbansound8K VocalSound Emotion Gender ESD Emotion Age Language TAU Urban Acoustic Scenes 2022 CochlScene BirdCLEF_2021 EmoBox AudioSet We also converted huge WAV files into MP3 16k sample rate to reduce storage size.To prevent leakage, please do not include test set in training session.… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Zeroshot-Audio-Classification-Instructions.audio1M<n<10M3 likes604 downloads1y agoHugging Face02jaeyong2 /cartesia-sonic-preview-ztts1-zero-shot-sample Cartesia Sonic on ZTTS1 zero-shot — sample with reference audio 100 utterances per language (700 rows) from the zero-shot subsets of ZTTS1-Eval, synthesized with Cartesia Sonic (preview) in voice-cloning mode. Unlike the full set, every row carries the reference recording as well as the synthesized clip, so a take can be compared against the voice it was cloning without checking out the benchmark. Columns column meaning audio the clip the model produced… See the full description on the dataset page: https://huggingface.co/datasets/jaeyong2/cartesia-sonic-preview-ztts1-zero-shot-sample.audiotext-to-speechn<1K0 likes48 downloads21d agoHugging Face03jaeyong2 /cartesia-sonic-preview-ztts1-zero-shot Cartesia sonic-preview — ZTTS1-Eval zero-shot synthesis + scores Zero-shot voice-cloning TTS on the ZTTS1-Eval benchmark (Zyphra, FLEURS-R based), now at full scale: 7 languages x 500 utterances = 3,500 per engine, five engines side by side: engine mode model Cartesia zero-shot clone sonic-preview (Sonic-3.6 beta at generation time) ElevenLabs zero-shot clone (IVC) eleven_v3 Qwen3-TTS Base zero-shot clone Qwen/Qwen3-TTS-12Hz-1.7B-Base (self-hosted, vLLM-Omni)… See the full description on the dataset page: https://huggingface.co/datasets/jaeyong2/cartesia-sonic-preview-ztts1-zero-shot.audio10K<n<100K0 likes13 downloads28d agoHugging Face04Finalprojectfour /test_zero_shot_dataaudion<1K0 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.