datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
massive-templates
Purpose. This dataset was collected specifically for intent-parser benchmarking, independently from any OVOS skill. Skill-derived utterances tend to overfit the exact phrasings a plugin was tuned on; this data is drawn from a disjoint source so it measures whether an OVOS intent plugin generalizes rather than memorizes. It is part of the OVOS intent-classification datasets used by the OVOS Plugin Arena intent benchmark.
Funding
Developed by TigreGotico for OpenVoiceOS as part… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/massive-templates.wake_word_noiseMT-intents-dataset-pt-PTovos-stt-bench-voxpopuli-en-US
OVOS stt bench — voxpopuli-en-US
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-en-US.ovos-stt-bench-mtedx-es-ES
OVOS stt bench — mtedx-es-ES
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
deepdml/mtedx.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mtedx-es-ES.ovos-community-wakewords-datasetMirror from https://github.com/OpenVoiceOS/ovos-ww-community-dataset
ovos-stt-bench-mls-pt-PT
OVOS stt bench — mls-pt-PT
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-pt-PT.ovos-stt-bench-mls-es-ES
OVOS stt bench — mls-es-ES
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-es-ES.ovos-stt-bench-ami-en-GB
OVOS stt bench — ami-en-GB
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
edinburghcstr/ami.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-ami-en-GB.ovos-stt-bench-speech-massive-fr-FR
OVOS stt bench — speech-massive-fr-FR
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-fr-FR.ovos-intent-bench-intents-for-eval
OVOS intent bench — intents-for-eval
Per-sample predictions of the open intent league (mixed-paradigm pipeline fusions) fighters of the
OVOS Plugin Arena over
OpenVoiceOS/intents-for-eval.
One dedicated repo per benchmark modality; one dataset split per language;
one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl.
Rows follow the arena §3.2 contract (pinned dataset_revision,
plugin_version, fired pipeline stage, exact_match with correct-OOD
semantics).… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-intents-for-eval.ovos-stt-bench-mtedx-de-DE
OVOS stt bench — mtedx-de-DE
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
deepdml/mtedx.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mtedx-de-DE.stt-sampler-v1
stt-sampler-v1
Licensing: clips inherit their source dataset's license — CC-BY-4.0
for MInDS-14 and FLEURS clips, CC-BY-NC-SA-4.0 for Speech-MASSIVE
clips (source_dataset column identifies each clip's origin).
A small, balanced, representative multilingual ASR eval sampler for the
OVOS Plugin Arena:
100 clips per language x 20 locales = 2000 clips, 16 kHz mono float32,
one config per locale (load_dataset("OpenVoiceOS/stt-sampler-v1", "<lang>")).
Designed to seed every STT… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/stt-sampler-v1.ovos-stt-bench-speech-massive-de-DE
OVOS stt bench — speech-massive-de-DE
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-de-DE.ovos-stt-bench-voxpopuli-es-ES
OVOS stt bench — voxpopuli-es-ES
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-es-ES.intents-for-eval
Purpose. This dataset was collected specifically for intent-parser benchmarking, independently from any OVOS skill. Skill-derived utterances tend to overfit the exact phrasings a plugin was tuned on; this data is drawn from a disjoint source so it measures whether an OVOS intent plugin generalizes rather than memorizes. It is part of the OVOS intent-classification datasets used by the OVOS Plugin Arena intent benchmark.
Funding
Developed by TigreGotico for OpenVoiceOS as part… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/intents-for-eval.ovos-stt-bench-mls-fr-FR
OVOS stt bench — mls-fr-FR
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-fr-FR.ovos-stt-bench-voxpopuli-fi-FI
OVOS stt bench — voxpopuli-fi-FI
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-fi-FI.ovos-stt-bench-voxpopuli-it-IT
OVOS stt bench — voxpopuli-it-IT
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-it-IT.ovos-stt-bench-mtedx-fr-FR
OVOS stt bench — mtedx-fr-FR
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
deepdml/mtedx.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mtedx-fr-FR.ovos-stt-bench-voxpopuli-sl-SI
OVOS stt bench — voxpopuli-sl-SI
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-sl-SI.ovos-stt-bench-mls-de-DE
OVOS stt bench — mls-de-DE
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-de-DE.ovos-stt-bench-voxpopuli-ro-RO
OVOS stt bench — voxpopuli-ro-RO
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-ro-RO.ovos-stt-bench-mtedx-it-IT
OVOS stt bench — mtedx-it-IT
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
deepdml/mtedx.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mtedx-it-IT.ovos-stt-bench-mls-it-IT
OVOS stt bench — mls-it-IT
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-it-IT.ovos-stt-bench-minds14-nl-NL
OVOS stt bench — minds14-nl-NL
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-minds14-nl-NL.ovos-intent-keyword-bench-intents-for-eval
OVOS intent_keyword bench — intents-for-eval
Per-sample predictions of the keyword-paradigm intent league fighters of the
OVOS Plugin Arena over
OpenVoiceOS/intents-for-eval.
One dedicated repo per benchmark modality; one dataset split per language;
one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl.
Rows follow the arena §3.2 contract (pinned dataset_revision,
plugin_version, fired pipeline stage, exact_match with correct-OOD
semantics). Produced by the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-keyword-bench-intents-for-eval.ovos-stt-bench-speech-massive-nl-NL
OVOS stt bench — speech-massive-nl-NL
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-nl-NL.ovos-wake-word-bench-community-ey-ordenador
OVOS wake_word bench — community-ey-ordenador
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/ovos-community-wakewords-dataset.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-ey-ordenador.ovos-wake-word-bench-community-hey-savant
OVOS wake_word bench — community-hey-savant
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/ovos-community-wakewords-dataset.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-hey-savant.
