datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voxpopuli
Dataset Card for Voxpopuli
Dataset Summary
VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.
The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials.
This implementation contains transcribed speech data for 18 languages.
It also contains 29 hours of transcribed speech data of non-native… See the full description on the dataset page: https://huggingface.co/datasets/facebook/voxpopuli.VoxPopuli-Cleaned-AA
VoxPopuli-Cleaned-AA
Quick links: AA Speech to Text Leaderboard | AA-WER v2.0 article
VoxPopuli-Cleaned-AA is a cleaned subset of the English VoxPopuli test data from esb/datasets, a speech dataset derived from European Parliament recordings. This cleaned subset is the VoxPopuli portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation of Speech to Text (STT) models.
This dataset is part of AA-WER… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/VoxPopuli-Cleaned-AA.voxpopuliA large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.voxpopuli
Dataset Card for Voxpopuli
Dataset Summary
VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.
The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials.
This implementation contains transcribed speech data for 18 languages.
It also contains 29 hours of transcribed speech data of non-native English… See the full description on the dataset page: https://huggingface.co/datasets/Prabh139/voxpopuli.VoxPopuli-Platinum-en-full
VoxPopuli-Platinum-en-full
Complete 142,066-row / ~404-hour English VoxPopuli Platinum dataset.
Reach out to data@trelis.com to purchase access or discuss a larger
custom-curation engagement.
Training Results
These results show why the Platinum labels matter. We compare the base model,
fine-tuning on raw VoxPopuli transcripts, and fine-tuning on this internally
filtered Platinum dataset. Evaluation uses the same english-spoken corpus WER
setup across four… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/VoxPopuli-Platinum-en-full.voxpopuli_es-ja
Dataset Card for Spanish-to-Japanese Automatic Speech Recognition Dataset
Dataset Summary
This dataset was created as part of a workshop organized by Yasmin Moslem, focusing on speech-to-text pipelines.
The workshop's primary goal is to enable accurate transcription and translation of spoken a source language into a written language (and learn how to do so, of course 😃)
The dataset serves as the foundation for developing and evaluating various models, including… See the full description on the dataset page: https://huggingface.co/datasets/marianogonzalezgomez/voxpopuli_es-ja.ellipsis-voxpopuli-ep-mouth-roi
Ellipsis — European Parliament Mouth-ROI (multilingual VSR)
Mouth-region-of-interest (ROI) clips + verbatim transcripts for visual speech recognition (VSR /
lip-reading), derived from European Parliament plenary debates (2019–2020) through the CC0
VoxPopuli segment index.
57,735 clips across 16 languages, in 58 WebDataset shards (~106 GB).
Each clip is the speaker's mouth ROI only — a [T, 96, 96] grayscale tensor (Auto-AVSR
convention) — paired with the VoxPopuli verbatim… See the full description on the dataset page: https://huggingface.co/datasets/TheNHz/ellipsis-voxpopuli-ep-mouth-roi.voxpopuliA large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.VoxPopuli-Platinum-en
VoxPopuli-Platinum-en
Showcase page and playable sample rows for the full 142,066-row / ~404-hour English VoxPopuli Platinum dataset.
Reach out to data@trelis.com to purchase access or discuss a larger
custom-curation engagement.
Training Results
These results show why the Platinum labels matter. We compare the base model,
fine-tuning on raw VoxPopuli transcripts, and fine-tuning on this internally
filtered Platinum dataset. Evaluation uses the same… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/VoxPopuli-Platinum-en.stt-voxpopuli-test-fr
VoxPopuli — French test split (mirror of facebook/voxpopuli)
Mirror public du split test de VoxPopuli config fr (Wang et al. ACL
2021, Facebook AI), pour benchmark ASR français parlementaire (sessions du
Parlement européen, 2009-2020, locuteurs MEP variés).
Ce repo ne contient que le split test FR (1 742 utterances ≈ 4-5 h).
Pour les splits train / validation ou les autres langues, voir le repo
upstream facebook/voxpopuli.
Contenu
1 742 utterances FR (eurodéputés… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-voxpopuli-test-fr.voxpopuli_es-ja
Dataset Card for Spanish-to-Japanese Automatic Speech Recognition Dataset
Dataset Summary
This dataset was created as part of a workshop organized by Yasmin Moslem, focusing on speech-to-text pipelines.
The workshop's primary goal is to enable accurate transcription and translation of spoken a source language into a written language (and learn how to do so, of course 😃)
The dataset serves as the foundation for developing and evaluating various models, including… See the full description on the dataset page: https://huggingface.co/datasets/Philou134/voxpopuli_es-ja.voxpopuli-timestampedA large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.stt-voxpopuli-test-en
VoxPopuli EN — test split
Mirror byte-exact du split test de facebook/voxpopuli config en.
1842 utt parlementaires (Parlement européen, 2009–2020)
16 kHz mono WAV embedded
Référence WER : normalized_text (préférée à raw_text)
Champ is_gold_transcript : True ⇒ transcription validée humaine
Source : Wang, Riviere et al. VoxPopuli: A Large-Scale Multilingual Speech
Corpus for Representation Learning, Semi-Supervised Learning and
Interpretation, ACL 2021.
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-voxpopuli-test-en.
