datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ovos-wake-word-bench-picovoice-snowboy
OVOS wake_word bench — picovoice-snowboy
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-snowboy.snowball-replay
Snowball Hugging Face row replay
This repository identifies the upstream Hugging Face rows retained in the Snowball pretraining store. It contains row locators and mixture metadata, not source documents or token arrays. A reader does not need Marin, private GCS access, or either of the earlier Snowball index repositories.
Get the selected rows
Install Python 3.12 or newer, then run:
python -m pip install huggingface_hub pyarrow zstandard
hf download… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/snowball-replay.
