wake-word
synthetic-wakewordslivekit_wakeword_featuresThis dataset contains precomputed audio features designed for use with the openWakeWord library.
Specifically, they are intended to be used as general purpose negative data (that is, data that does not contain the target wake word/phrase) for training custom openWakeWord models.
The individual .npy files in this dataset are not original audio data, but rather are low dimensional audio features produced by a pre-trained speech embedding model from Google.
openWakeWord uses these features as… See the full description on the dataset page: https://huggingface.co/datasets/binhpham/livekit_wakeword_features.sam-wake-word-raw-datanot-wake-words-speech-en
not-wake-words-speech-en
Negative (non-wake-word) speech clips, used to measure false accepts for OVOS
wake-word plugins.
Derived from the Multilingual Spoken Words Corpus
(MLCommons), which is built from Mozilla Common Voice and licensed CC-BY-4.0.
This derivative keeps the same licence and attribution requirement.
Produced with support from the NGI0 Commons Fund.
synthetic-wakeword-hey_computer
synthetic-wakeword-hey_computer
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "hey computer".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_computer.synthetic-wakewords
synthetic-wakewords
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "multiple wake words".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on, including… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/synthetic-wakewords.
