openvoiceos
ovos-tts-bench-massive-prompts
OVOS tts bench — massive-prompts
Synthesised clips (one per prompt) predictions of the registered
OVOS Plugin Arena
tts fighters over
OpenVoiceOS/massive-templates.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-tts-bench-massive-prompts.ovos-stt-bench-stt-sampler-v1
OVOS stt bench — stt-sampler-v1
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
OpenVoiceOS/stt-sampler-v1.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-stt-sampler-v1.ovos-tts-bench-intents-for-eval-prompts
OVOS tts bench — intents-for-eval-prompts
Synthesised clips (one per prompt) predictions of the registered
OVOS Plugin Arena
tts fighters over
OpenVoiceOS/intents-for-eval.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-tts-bench-intents-for-eval-prompts.massive-templates
Purpose. This dataset was collected specifically for intent-parser benchmarking, independently from any OVOS skill. Skill-derived utterances tend to overfit the exact phrasings a plugin was tuned on; this data is drawn from a disjoint source so it measures whether an OVOS intent plugin generalizes rather than memorizes. It is part of the OVOS intent-classification datasets used by the OVOS Plugin Arena intent benchmark.
Funding
Developed by TigreGotico for OpenVoiceOS as part… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/massive-templates.ovos-localize-intents
OpenVoiceOS Localize — Intent Classification Dataset
Multilingual intent classification corpus exported from
OpenVoiceOS/ovos-localize.
Each row is a single expanded utterance labelled with the OVOS skill and
intent file that produced it. Templates are fully expanded (bracket
alternation resolved); {slot_name} placeholders from .intent files are
kept verbatim so models can learn the slot-carrying pattern.
Schema
Column
Description
lang
BCP-47 locale… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-localize-intents.synthetic-wakewords
synthetic-wakewords
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "multiple wake words".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on, including… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/synthetic-wakewords.
