datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ovos-vad-bench-speech-vs-nonspeech-cs-CZ
OVOS vad bench — speech-vs-nonspeech-cs-CZ
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-cs-CZ.ovos-vad-bench-speech-vs-nonspeech-en-US
OVOS vad bench — speech-vs-nonspeech-en-US
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-en-US.ovos-vad-bench-speech-vs-nonspeech-en-GB
OVOS vad bench — speech-vs-nonspeech-en-GB
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-en-GB.ovos-vad-bench-speech-vs-nonspeech-en-AU
OVOS vad bench — speech-vs-nonspeech-en-AU
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-en-AU.ovos-vad-bench-speech-vs-nonspeech-es-ES
OVOS vad bench — speech-vs-nonspeech-es-ES
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-es-ES.ovos-vad-bench-speech-vs-nonspeech-de-DE
OVOS vad bench — speech-vs-nonspeech-de-DE
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-de-DE.ovos-vad-bench-speech-vs-nonspeech-hu-HU
OVOS vad bench — speech-vs-nonspeech-hu-HU
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-hu-HU.ovos-vad-bench-speech-vs-nonspeech-pl-PL
OVOS vad bench — speech-vs-nonspeech-pl-PL
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-pl-PL.ovos-vad-bench-speech-vs-nonspeech-tr-TR
OVOS vad bench — speech-vs-nonspeech-tr-TR
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-tr-TR.ovos-vad-bench-speech-vs-nonspeech-it-IT
OVOS vad bench — speech-vs-nonspeech-it-IT
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-it-IT.ovos-vad-bench-speech-vs-nonspeech-pt-PT
OVOS vad bench — speech-vs-nonspeech-pt-PT
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-pt-PT.ovos-vad-bench-speech-vs-nonspeech-vi-VN
OVOS vad bench — speech-vs-nonspeech-vi-VN
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-vi-VN.ovos-vad-bench-speech-vs-nonspeech-zh-CN
OVOS vad bench — speech-vs-nonspeech-zh-CN
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-zh-CN.ovos-vad-bench-speech-vs-nonspeech-fr-FR
OVOS vad bench — speech-vs-nonspeech-fr-FR
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-fr-FR.ovos-vad-bench-speech-vs-nonspeech-nl-NL
OVOS vad bench — speech-vs-nonspeech-nl-NL
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-nl-NL.ovos-vad-bench-speech-vs-nonspeech-ko-KR
OVOS vad bench — speech-vs-nonspeech-ko-KR
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-ko-KR.ovos-vad-bench-speech-vs-nonspeech-ru-RU
OVOS vad bench — speech-vs-nonspeech-ru-RU
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-ru-RU.hi-en-noisy-vad-benchmark
Hindi-English Noisy VAD Benchmark
Version 0.1.0 is a deterministic, evaluation-only benchmark with 78
mono PCM16 WAV files at 16 kHz: six clean speech controls and 72 mixtures spanning
six speech sources, three real noise categories, and four SNRs (20, 10, 5, 0 dB).
Intended use
Use this dataset to compare voice-activity detectors under matched Hindi/English
noise conditions and to tune thresholds. It is too small and insufficiently diverse
for model training… See the full description on the dataset page: https://huggingface.co/datasets/Aakash22134/hi-en-noisy-vad-benchmark.findfind2
