edgeimpulse/Hey-Edge
hey_edge — Wake Word Synthetic Speech Dataset Synthetic, augmented audio for training a small wake-word / keyword-spotting model. Generated with piper_tts and local audio augmentation. Classes Label Samples background_noise 200 hey_edge 378 unknown 1071 hey_edge — the target wake phrase and close variants. unknown — near-miss and unrelated short phrases. background_noise — synthetic background noise. Audio Specification… See the full description on the dataset page: https://huggingface.co/datasets/edgeimpulse/Hey-Edge.
hey_edge — Wake Word Synthetic Speech Dataset
Synthetic, augmented audio for training a small wake-word / keyword-spotting model. Generated with piper_tts and local audio augmentation.
Classes
hey_edge— the target wake phrase and close variants.unknown— near-miss and unrelated short phrases.background_noise— synthetic background noise.
Audio Specification
Layout
Audio is organised into one sub-folder per class so the Hugging Face dataset viewer infers the label column automatically:
audio/
train/
hey_edge/ hey_edge.<id>.wav ...
unknown/ unknown.<id>.wav ...
background_noise/ background_noise.<id>.wav ...
test/
hey_edge/ ...
unknown/ ...
background_noise/ ...
edge_impulse_metadata.csv
hf_metadata.csv
selected_voices.csv
dataset_summary.jsonLoading
from datasets import load_dataset, Audio
ds = load_dataset("edgeimpulse/Hey-Edge")
ds = ds.cast_column("audio", Audio(sampling_rate=16000))
print(ds)Edge Impulse
Filenames follow the Edge Impulse label-prefix convention (hey_edge.<id>.wav) so they upload directly:
edge-impulse-uploader --category training audio/train/**/*.wav
edge-impulse-uploader --category testing audio/test/**/*.wavLimitations
Synthetic TTS is a bootstrap, not a production benchmark. Add real device and environment recordings before deploying a wake-word product.
License
CC BY 4.0. Verify that your use of the generated synthetic speech complies with the terms of the voice models and tools used to create it.
