CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yojul /one-shot-hip-hop-drumsaudio10K<n<100K4 likes77 downloads2y agoHugging Face02danielrosehill /Whisper-Fine-Tune-One-Shot-Eval Whisper Fine-Tuning Evaluation: Local vs Commercial ASR A "back of the envelope" evaluation comparing fine-tuned Whisper models running locally against commercial ASR APIs via Eden AI. The Question Can fine-tuning Whisper achieve measurable WER reductions, even when comparing local inference against cloud-based commercial models? TL;DR Yes. Fine-tuned Whisper Large Turbo running locally achieved 5.84% WER, beating the best commercial API (Assembly at… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Whisper-Fine-Tune-One-Shot-Eval.audion<1K0 likes53 downloads10mo agoHugging Face03jaeyong2 /cartesia-sonic-preview-ztts1-zero-shot-sample Cartesia Sonic on ZTTS1 zero-shot — sample with reference audio 100 utterances per language (700 rows) from the zero-shot subsets of ZTTS1-Eval, synthesized with Cartesia Sonic (preview) in voice-cloning mode. Unlike the full set, every row carries the reference recording as well as the synthesized clip, so a take can be compared against the voice it was cloning without checking out the benchmark. Columns column meaning audio the clip the model produced… See the full description on the dataset page: https://huggingface.co/datasets/jaeyong2/cartesia-sonic-preview-ztts1-zero-shot-sample.audiotext-to-speechn<1K0 likes50 downloads23d agoHugging Face04mteb /Shot2Story20K_testaudio1K<n<10K0 likes15 downloads7mo agoHugging Face05fillipean /shotokanaudion<1K0 likes11 downloads3y agoHugging Face06Finalprojectfour /test_zero_shot_dataaudion<1K0 likes9 downloads2y agoHugging Face07jaeyong2 /cartesia-sonic-preview-ztts1-zero-shot Cartesia sonic-preview — ZTTS1-Eval zero-shot synthesis + scores Zero-shot voice-cloning TTS on the ZTTS1-Eval benchmark (Zyphra, FLEURS-R based), now at full scale: 7 languages x 500 utterances = 3,500 per engine, five engines side by side: engine mode model Cartesia zero-shot clone sonic-preview (Sonic-3.6 beta at generation time) ElevenLabs zero-shot clone (IVC) eleven_v3 Qwen3-TTS Base zero-shot clone Qwen/Qwen3-TTS-12Hz-1.7B-Base (self-hosted, vLLM-Omni)… See the full description on the dataset page: https://huggingface.co/datasets/jaeyong2/cartesia-sonic-preview-ztts1-zero-shot.audio10K<n<100K0 likes5 downloads1mo agoHugging Face08shotasun /shotakun-kikkakeaudion<1K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.