CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01igidn /wuwa-voice-EN wuwa-voice-EN wuwa voice EN is a dataset of voice line from Wuthering Waves Attribute Value Language English Total Samples 37,805 Total Duration ~24 GB (WAV format) Unique Speakers 913 Categories 106 Audio Format WAV Transcription Format Plain text Field Description file_name Relative path to the audio file (e.g., data/en_vo_Category_1_1.wav) text transcript speaker Character (e.g., Zani, Carlotta, {PlayerName}) speaker_id Numeric ID… See the full description on the dataset page: https://huggingface.co/datasets/igidn/wuwa-voice-EN.audiotext-to-speech10K<n<100K2 likes627 downloads4mo agoHugging Face02bhyuan /wuw_datasetgated Yougen/wuw_dataset Wake-Up-Word (WUW) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where multiple utterances share a long recording via the segments file. To avoid duplicating audio, each tar sample corresponds to one full recording. The utterance-level metadata (id / start / end / text / spk / duration) is stored in a JSON list inside that sample. Downstream consumers slice the decoded… See the full description on the dataset page: https://huggingface.co/datasets/bhyuan/wuw_dataset.audioautomatic-speech-recognition0 likes6 downloads2mo agoHugging Face03ygyuan /wuw_mingated ygyuan/wuw_min Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: train: 65 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw audio… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/wuw_min.audioaudio-classification0 likes6 downloads2mo agoHugging Face04heimayuan /wuw_cszhengated ygyuan/wuw_cszhen Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: train: 712 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw audio… See the full description on the dataset page: https://huggingface.co/datasets/heimayuan/wuw_cszhen.audioaudio-classification0 likes6 downloads2mo agoHugging Face05Yougen /wuw_accentgated ygyuan/wuw_accent Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: train: 509 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw audio… See the full description on the dataset page: https://huggingface.co/datasets/Yougen/wuw_accent.audioaudio-classification0 likes5 downloads2mo agoHugging Face06ygyuan /wuw_wuygated ygyuan/wuw_wuy Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: train: 99 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw audio… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/wuw_wuy.audioaudio-classification0 likes4 downloads2mo agoHugging Face07heimayuan /wuw_testset1gated Yougen/wuw_testset1 Wake-Up-Word (WUW) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where multiple utterances share a long recording via the segments file. To avoid duplicating audio, each tar sample corresponds to one full recording. The utterance-level metadata (id / start / end / text / spk / duration) is stored in a JSON list inside that sample. Downstream consumers slice the decoded… See the full description on the dataset page: https://huggingface.co/datasets/heimayuan/wuw_testset1.audioautomatic-speech-recognition0 likes4 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.