datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voxpopuli
Dataset Card for Voxpopuli
Dataset Summary
VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.
The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials.
This implementation contains transcribed speech data for 18 languages.
It also contains 29 hours of transcribed speech data of non-native… See the full description on the dataset page: https://huggingface.co/datasets/facebook/voxpopuli.voxpopuli_mosel_curatorvoxpopuli
Dataset Card for "voxpopuli"
More Information needed
VoxPopuli-Cleaned-AA
VoxPopuli-Cleaned-AA
Quick links: AA Speech to Text Leaderboard | AA-WER v2.0 article
VoxPopuli-Cleaned-AA is a cleaned subset of the English VoxPopuli test data from esb/datasets, a speech dataset derived from European Parliament recordings. This cleaned subset is the VoxPopuli portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation of Speech to Text (STT) models.
This dataset is part of AA-WER… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/VoxPopuli-Cleaned-AA.VoxPopuliAccentPairClassificationFilteredvoxpopuli_elvoxpopuli-minimini-voxpopulivoxpopuli-mls-de-descriptions
Natural Language Voice Descriptions of the VoxPopuli and MLS German Datasets
German read and parliamentary speech paired with its transcript, acoustic
measurements, discrete German descriptor tags, and a free-text German
description of the speaker's voice and recording conditions. The dataset is intended
for training description-conditioned TTS models such as
Parler-TTS.
The data was built as part of research work. It is a random subset of the
pooled German portions of VoxPopuli… See the full description on the dataset page: https://huggingface.co/datasets/leonhard-behr/voxpopuli-mls-de-descriptions.voxpopuli-test-words-tempvoxpopuli_en_pseudo_labelledvoxpopuli-es-sr-24000voxpopuliA large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.ovos-stt-bench-voxpopuli-en-US
OVOS stt bench — voxpopuli-en-US
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-en-US.Voxpopuli_NER
VoxPopuli_NER
VoxPopuli-NER is derived from the VoxPopuli corpus and specifically enhanced for
Named Entity Recognition (NER) tasks focusing on political and geographical entities.
It includes 879 audio samples, annotated with 2469 unique entity types. The dataset consists of the English part of the test set of VoxPopuli.
See full details in the WhisperNER paper.
citation
If you find this usful, please cite the following works:
@article{ayache2024whisperner… See the full description on the dataset page: https://huggingface.co/datasets/aiola/Voxpopuli_NER.VoxPopuli-Cleaned-AA-with-audioSame as ArtificialAnalysis/VoxPopuli-Cleaned-AA but added back the audio from the original voxpopuli dataset
voxpopuli-fr-duration
Dataset Card for "voxpopuli-fr-duration"
More Information needed
voxpopuli_asr_curatorvoxpopuli_es_pseudo_labelledspanish_voxpopuli_alignedovos-stt-bench-voxpopuli-en-001
OVOS stt bench — voxpopuli-en-001
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-en-001.ovos-stt-bench-voxpopuli-es-ES
OVOS stt bench — voxpopuli-es-ES
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-es-ES.ovos-stt-bench-voxpopuli-it-IT
OVOS stt bench — voxpopuli-it-IT
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-it-IT.ovos-stt-bench-voxpopuli-fi-FI
OVOS stt bench — voxpopuli-fi-FI
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-fi-FI.ovos-stt-bench-voxpopuli-sl-SI
OVOS stt bench — voxpopuli-sl-SI
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-sl-SI.ovos-stt-bench-voxpopuli-ro-RO
OVOS stt bench — voxpopuli-ro-RO
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-ro-RO.ovos-stt-bench-voxpopuli-nl-NL
OVOS stt bench — voxpopuli-nl-NL
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-nl-NL.ovos-stt-bench-voxpopuli-fr-FR
OVOS stt bench — voxpopuli-fr-FR
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-fr-FR.ovos-stt-bench-voxpopuli-sk-SK
OVOS stt bench — voxpopuli-sk-SK
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-sk-SK.ovos-stt-bench-voxpopuli-de-DE
OVOS stt bench — voxpopuli-de-DE
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-de-DE.
