CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PoojasreeBalasubramanian /synthetic-wakewordsaudio10K<n<100K0 likes3k downloads2mo agoHugging Face02binhpham /livekit_wakeword_featuresThis dataset contains precomputed audio features designed for use with the openWakeWord library. Specifically, they are intended to be used as general purpose negative data (that is, data that does not contain the target wake word/phrase) for training custom openWakeWord models. The individual .npy files in this dataset are not original audio data, but rather are low dimensional audio features produced by a pre-trained speech embedding model from Google. openWakeWord uses these features as… See the full description on the dataset page: https://huggingface.co/datasets/binhpham/livekit_wakeword_features.0 likes3k downloads7mo agoHugging Face03quo-labs /sam-wake-word-raw-dataaudio10K<n<100K2 likes1.1k downloads23d agoHugging Face04TigreGotico /not-wake-words-speech-en not-wake-words-speech-en Negative (non-wake-word) speech clips, used to measure false accepts for OVOS wake-word plugins. Derived from the Multilingual Spoken Words Corpus (MLCommons), which is built from Mozilla Common Voice and licensed CC-BY-4.0. This derivative keeps the same licence and attribution requirement. Produced with support from the NGI0 Commons Fund. audioaudio-classification10K<n<100K0 likes584 downloads20d agoHugging Face05TigreGotico /synthetic-wakeword-hey_computer synthetic-wakeword-hey_computer Synthetic wake-word audio for training and benchmarking OVOS wake-word plugins, covering the phrase "hey computer". Every clip is machine-generated: text-to-speech synthesis followed by voice conversion to simulate multiple speakers. No human recording is included, and no natural voice is reproduced. Machine-generated audio carries no copyright of its own, so this dataset is published CC-BY-4.0 and is free to use, redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_computer.audioaudio-classification1K<n<10K0 likes405 downloads6d agoHugging Face06OpenVoiceOS /synthetic-wakewords synthetic-wakewords Synthetic wake-word audio for training and benchmarking OVOS wake-word plugins, covering the phrase "multiple wake words". Every clip is machine-generated: text-to-speech synthesis followed by voice conversion to simulate multiple speakers. No human recording is included, and no natural voice is reproduced. Machine-generated audio carries no copyright of its own, so this dataset is published CC-BY-4.0 and is free to use, redistribute and build on, including… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/synthetic-wakewords.audioaudio-classification10K<n<100K3 likes353 downloads20d agoHugging Face07OpenVoiceOS /wake_word_noiseaudio1K<n<10K0 likes322 downloads2y agoHugging Face08AigizK /homai_wake_word_1020ms Homai Wake Word 1020 ms Binary training corpus for a wake-word detector that should react to Һомай and Хомай, and reject other speech. Every audio value is mono 16 kHz FLAC with exactly 16,320 samples (1020 ms). Audited size Split Rows train 1,589,300 validation 87,701 test 87,776 Total 1,764,777 Label Rows positive 833,696 negative 931,081 Source rows in the release: AigizK/Homai-Wake-Word (5,486)… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/homai_wake_word_1020ms.audioaudio-classification1M<n<10M0 likes288 downloads2mo agoHugging Face09AigizK /homai_wake_word_omnivoice Homai Wake Word OmniVoice Synthetic two-label wake-word dataset generated with k2-fsa/OmniVoice using cross-lingual voice cloning. For every reference row from all train, validation, and test splits of: bond005/sova_rudevices bond005/sberdevices_golos_100h_farfield the dataset contains two generated recordings: Һомай, generated with OmniVoice language Bashkir; Хомай, generated with OmniVoice language Russian. Dataset structure Split: train Columns: audio… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/homai_wake_word_omnivoice.audioaudio-classification100K<n<1M0 likes285 downloads2mo agoHugging Face10OpenVoiceOS /ovos-wake-word-bench-synthetic-wakewords-hey_ziggy OVOS wake_word bench — synthetic-wakewords-hey_ziggy Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/synthetic-wakewords. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_ziggy.0 likes228 downloads16d agoHugging Face11OpenVoiceOS /ovos-community-wakewords-datasetMirror from https://github.com/OpenVoiceOS/ovos-ww-community-dataset audioaudio-classification1K<n<10K0 likes226 downloads2y agoHugging Face12OpenVoiceOS /ovos-wake-word-bench-synthetic-wakewords-hey_mycroft OVOS wake_word bench — synthetic-wakewords-hey_mycroft Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/synthetic-wakewords. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_mycroft.0 likes192 downloads16d agoHugging Face13OpenVoiceOS /ovos-wake-word-bench-community-ey-ordenador OVOS wake_word bench — community-ey-ordenador Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/ovos-community-wakewords-dataset. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-ey-ordenador.tabularn<1K0 likes175 downloads16d agoHugging Face14OpenVoiceOS /ovos-wake-word-bench-community-hey-savant OVOS wake_word bench — community-hey-savant Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/ovos-community-wakewords-dataset. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-hey-savant.tabular1K<n<10K0 likes170 downloads16d agoHugging Face15OpenVoiceOS /ovos-wake-word-bench-community-hey-floyd OVOS wake_word bench — community-hey-floyd Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/ovos-community-wakewords-dataset. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-hey-floyd.tabular1K<n<10K0 likes166 downloads16d agoHugging Face16OpenVoiceOS /ovos-wake-word-bench-community-computer OVOS wake_word bench — community-computer Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/ovos-community-wakewords-dataset. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-computer.tabularn<1K0 likes164 downloads16d agoHugging Face17OpenVoiceOS /ovos-wake-word-bench-community-hey-ziggy OVOS wake_word bench — community-hey-ziggy Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/ovos-community-wakewords-dataset. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-hey-ziggy.tabular1K<n<10K0 likes154 downloads16d agoHugging Face18OpenVoiceOS /ovos-wake-word-bench-picovoice-computer OVOS wake_word bench — picovoice-computer Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over Picovoice/wake-word-benchmark. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-computer.tabular1K<n<10K0 likes152 downloads16d agoHugging Face19TigreGotico /synthetic-wakeword-hey_mycroft synthetic-wakeword-hey_mycroft Synthetic wake-word audio for training and benchmarking OVOS wake-word plugins, covering the phrase "hey mycroft". Every clip is machine-generated: text-to-speech synthesis followed by voice conversion to simulate multiple speakers. No human recording is included, and no natural voice is reproduced. Machine-generated audio carries no copyright of its own, so this dataset is published CC-BY-4.0 and is free to use, redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_mycroft.audioaudio-classification1K<n<10K0 likes147 downloads6d agoHugging Face20OpenVoiceOS /ovos-wake-word-bench-picovoice-alexa OVOS wake_word bench — picovoice-alexa Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over Picovoice/wake-word-benchmark. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-alexa.tabular1K<n<10K0 likes147 downloads16d agoHugging Face21OpenVoiceOS /ovos-wake-word-bench-picovoice-smart-mirror OVOS wake_word bench — picovoice-smart-mirror Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over Picovoice/wake-word-benchmark. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-smart-mirror.tabular1K<n<10K0 likes144 downloads16d agoHugging Face22OpenVoiceOS /ovos-wake-word-bench-community-athena OVOS wake_word bench — community-athena Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/ovos-community-wakewords-dataset. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-athena.tabular1K<n<10K0 likes138 downloads16d agoHugging Face23Picovoice /wake-word-benchmark Wake Word Benchmark Made in Vancouver, Canada by Picovoice The purpose of this benchmarking framework is to provide a scientific comparison between different wake word detection engines in terms of accuracy and runtime metrics. While working on Porcupine we noted that there is a need for such a tool to empower customers to make data-driven decisions. Results Accuracy Below is the result of running the benchmark framework averaged on six different… See the full description on the dataset page: https://huggingface.co/datasets/Picovoice/wake-word-benchmark.audio1K<n<10K0 likes136 downloads10mo agoHugging Face24OpenVoiceOS /ovos-wake-word-bench-picovoice-jarvis OVOS wake_word bench — picovoice-jarvis Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over Picovoice/wake-word-benchmark. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-jarvis.tabular1K<n<10K0 likes136 downloads16d agoHugging Face25OpenVoiceOS /ovos-wake-word-bench-community-amelia OVOS wake_word bench — community-amelia Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/ovos-community-wakewords-dataset. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-amelia.0 likes132 downloads16d agoHugging Face26OpenVoiceOS /ovos-wake-word-bench-synthetic-wakewords-hey_jarvis OVOS wake_word bench — synthetic-wakewords-hey_jarvis Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/synthetic-wakewords. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_jarvis.tabularn<1K0 likes130 downloads16d agoHugging Face27OpenVoiceOS /ovos-wake-word-bench-picovoice-snowboy OVOS wake_word bench — picovoice-snowboy Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over Picovoice/wake-word-benchmark. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-snowboy.tabular1K<n<10K0 likes130 downloads16d agoHugging Face28OpenVoiceOS /ovos-wake-word-bench-mlsw-negatives-en-US OVOS wake_word bench — mlsw-negatives-en-US Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over MLCommons/ml_spoken_words. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-mlsw-negatives-en-US.0 likes128 downloads18d agoHugging Face29AlazarM /hey-native-wakeword Hey Native — wake-word dataset Synthetic 16 kHz mono audio for training a "Hey Native" wake-word detector. data/positive/ — utterances of "Hey Native" (label 1) data/negative/ — general speech, not the wake word (label 0) data/hard_negative/ — near-miss confusables, e.g. "hey navy", "hey maybe" (label 0) metadata.csv — file_name, label, label_name, text, source_model, mos_p808 5,000 positives · 6,000 negatives · 1,050 hard negatives. audioaudio-classification10K<n<100K0 likes123 downloads7d agoHugging Face30TigreGotico /synthetic-wakeword-hey_siri synthetic-wakeword-hey_siri Synthetic wake-word audio for training and benchmarking OVOS wake-word plugins, covering the phrase "hey siri". Every clip is machine-generated: text-to-speech synthesis followed by voice conversion to simulate multiple speakers. No human recording is included, and no natural voice is reproduced. Machine-generated audio carries no copyright of its own, so this dataset is published CC-BY-4.0 and is free to use, redistribute and build on, including for… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_siri.audioaudio-classificationn<1K0 likes119 downloads20d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.