ygyuan/kws_testset_ug_huiting
ygyuan/kws_testset_ug_huiting Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav #… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ug_huiting.
This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.
