CoolFace
Datasetpublicgated

heimayuan/wuw_testset1

Yougen/wuw_testset1 Wake-Up-Word (WUW) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where multiple utterances share a long recording via the segments file. To avoid duplicating audio, each tar sample corresponds to one full recording. The utterance-level metadata (id / start / end / text / spk / duration) is stored in a JSON list inside that sample. Downstream consumers slice the decoded… See the full description on the dataset page: https://huggingface.co/datasets/heimayuan/wuw_testset1.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes4downloads
settings

This repository belongs to heimayuan on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namewuw_testset1
visibilitypublic
licenceother
gatedyes
ownerheimayuan
Account settings
heimayuan/wuw_testset1 · CoolFace