datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-wakeword-hey_computer
synthetic-wakeword-hey_computer
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "hey computer".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_computer.hey-computer-speech-commands
Hey Computer: Speech Command Recognition
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
13,192
Labeled training data
test
3,295
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
audio
Audio
id
string
label
string (test sentinel: unlabeled)… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/hey-computer-speech-commands.badini-tts-samplescommonvoice-kmrsynthetic-wakeword-hey_computer
