datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-wakewordslivekit_wakeword_featuresThis dataset contains precomputed audio features designed for use with the openWakeWord library.
Specifically, they are intended to be used as general purpose negative data (that is, data that does not contain the target wake word/phrase) for training custom openWakeWord models.
The individual .npy files in this dataset are not original audio data, but rather are low dimensional audio features produced by a pre-trained speech embedding model from Google.
openWakeWord uses these features as… See the full description on the dataset page: https://huggingface.co/datasets/binhpham/livekit_wakeword_features.sam-wake-word-raw-datanot-wake-words-speech-en
not-wake-words-speech-en
Negative (non-wake-word) speech clips, used to measure false accepts for OVOS
wake-word plugins.
Derived from the Multilingual Spoken Words Corpus
(MLCommons), which is built from Mozilla Common Voice and licensed CC-BY-4.0.
This derivative keeps the same licence and attribution requirement.
Produced with support from the NGI0 Commons Fund.
synthetic-wakeword-hey_computer
synthetic-wakeword-hey_computer
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "hey computer".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_computer.synthetic-wakewords
synthetic-wakewords
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "multiple wake words".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on, including… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/synthetic-wakewords.wake_word_noisehomai_wake_word_1020ms
Homai Wake Word 1020 ms
Binary training corpus for a wake-word detector that should react to Һомай
and Хомай, and reject other speech. Every audio value is mono 16 kHz FLAC
with exactly 16,320 samples (1020 ms).
Audited size
Split
Rows
train
1,589,300
validation
87,701
test
87,776
Total
1,764,777
Label
Rows
positive
833,696
negative
931,081
Source rows in the release: AigizK/Homai-Wake-Word
(5,486)… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/homai_wake_word_1020ms.homai_wake_word_omnivoice
Homai Wake Word OmniVoice
Synthetic two-label wake-word dataset generated with k2-fsa/OmniVoice using
cross-lingual voice cloning.
For every reference row from all train, validation, and test splits of:
bond005/sova_rudevices
bond005/sberdevices_golos_100h_farfield
the dataset contains two generated recordings:
Һомай, generated with OmniVoice language Bashkir;
Хомай, generated with OmniVoice language Russian.
Dataset structure
Split: train
Columns: audio… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/homai_wake_word_omnivoice.ovos-wake-word-bench-synthetic-wakewords-hey_ziggy
OVOS wake_word bench — synthetic-wakewords-hey_ziggy
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_ziggy.ovos-community-wakewords-datasetMirror from https://github.com/OpenVoiceOS/ovos-ww-community-dataset
ovos-wake-word-bench-synthetic-wakewords-hey_mycroft
OVOS wake_word bench — synthetic-wakewords-hey_mycroft
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_mycroft.ovos-wake-word-bench-community-ey-ordenador
OVOS wake_word bench — community-ey-ordenador
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/ovos-community-wakewords-dataset.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-ey-ordenador.ovos-wake-word-bench-community-hey-savant
OVOS wake_word bench — community-hey-savant
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/ovos-community-wakewords-dataset.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-hey-savant.ovos-wake-word-bench-community-hey-floyd
OVOS wake_word bench — community-hey-floyd
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/ovos-community-wakewords-dataset.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-hey-floyd.ovos-wake-word-bench-community-computer
OVOS wake_word bench — community-computer
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/ovos-community-wakewords-dataset.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-computer.ovos-wake-word-bench-community-hey-ziggy
OVOS wake_word bench — community-hey-ziggy
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/ovos-community-wakewords-dataset.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-hey-ziggy.ovos-wake-word-bench-picovoice-computer
OVOS wake_word bench — picovoice-computer
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-computer.synthetic-wakeword-hey_mycroft
synthetic-wakeword-hey_mycroft
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "hey mycroft".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_mycroft.ovos-wake-word-bench-picovoice-alexa
OVOS wake_word bench — picovoice-alexa
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-alexa.ovos-wake-word-bench-picovoice-smart-mirror
OVOS wake_word bench — picovoice-smart-mirror
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-smart-mirror.ovos-wake-word-bench-community-athena
OVOS wake_word bench — community-athena
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/ovos-community-wakewords-dataset.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-athena.wake-word-benchmark
Wake Word Benchmark
Made in Vancouver, Canada by Picovoice
The purpose of this benchmarking framework is to provide a scientific comparison between different wake word detection
engines in terms of accuracy and runtime metrics. While working on Porcupine
we noted that there is a need for such a tool to empower customers to make data-driven decisions.
Results
Accuracy
Below is the result of running the benchmark framework averaged on six different… See the full description on the dataset page: https://huggingface.co/datasets/Picovoice/wake-word-benchmark.ovos-wake-word-bench-picovoice-jarvis
OVOS wake_word bench — picovoice-jarvis
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-jarvis.ovos-wake-word-bench-community-amelia
OVOS wake_word bench — community-amelia
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/ovos-community-wakewords-dataset.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-amelia.ovos-wake-word-bench-synthetic-wakewords-hey_jarvis
OVOS wake_word bench — synthetic-wakewords-hey_jarvis
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_jarvis.ovos-wake-word-bench-picovoice-snowboy
OVOS wake_word bench — picovoice-snowboy
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-snowboy.ovos-wake-word-bench-mlsw-negatives-en-US
OVOS wake_word bench — mlsw-negatives-en-US
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
MLCommons/ml_spoken_words.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-mlsw-negatives-en-US.hey-native-wakeword
Hey Native — wake-word dataset
Synthetic 16 kHz mono audio for training a "Hey Native" wake-word detector.
data/positive/ — utterances of "Hey Native" (label 1)
data/negative/ — general speech, not the wake word (label 0)
data/hard_negative/ — near-miss confusables, e.g. "hey navy", "hey maybe" (label 0)
metadata.csv — file_name, label, label_name, text, source_model, mos_p808
5,000 positives · 6,000 negatives · 1,050 hard negatives.
synthetic-wakeword-hey_siri
synthetic-wakeword-hey_siri
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "hey siri".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on, including for… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_siri.
