datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voxpopuli
Dataset Card for Voxpopuli
Dataset Summary
VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.
The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials.
This implementation contains transcribed speech data for 18 languages.
It also contains 29 hours of transcribed speech data of non-native… See the full description on the dataset page: https://huggingface.co/datasets/facebook/voxpopuli.voxpopuli
Dataset Card for "voxpopuli"
More Information needed
VoxPopuli-Cleaned-AA
VoxPopuli-Cleaned-AA
Quick links: AA Speech to Text Leaderboard | AA-WER v2.0 article
VoxPopuli-Cleaned-AA is a cleaned subset of the English VoxPopuli test data from esb/datasets, a speech dataset derived from European Parliament recordings. This cleaned subset is the VoxPopuli portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation of Speech to Text (STT) models.
This dataset is part of AA-WER… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/VoxPopuli-Cleaned-AA.VoxPopuliAccentPairClassificationFilteredvoxpopuli_elvoxpopuli-minimini-voxpopulivoxpopuli-mls-de-descriptions
Natural Language Voice Descriptions of the VoxPopuli and MLS German Datasets
German read and parliamentary speech paired with its transcript, acoustic
measurements, discrete German descriptor tags, and a free-text German
description of the speaker's voice and recording conditions. The dataset is intended
for training description-conditioned TTS models such as
Parler-TTS.
The data was built as part of research work. It is a random subset of the
pooled German portions of VoxPopuli… See the full description on the dataset page: https://huggingface.co/datasets/leonhard-behr/voxpopuli-mls-de-descriptions.voxpopuli-es-sr-24000voxpopuli_en_pseudo_labelledvoxpopuli-fr-duration
Dataset Card for "voxpopuli-fr-duration"
More Information needed
voxpopuli_es_pseudo_labelledVoxPopuli-Cleaned-AA-with-audioSame as ArtificialAnalysis/VoxPopuli-Cleaned-AA but added back the audio from the original voxpopuli dataset
spanish_voxpopuli_alignedvoxpopuliit-voxpopuli-pseudo-labeled-whisper-large-v3voxpopuli-qc-samples-v3
VoxPopuli QC Samples V3 - CER-based Quality Control
Quality control samples from VoxPopuli French ASR pseudolabeling, categorized by Character Error Rate (CER) between Whisper (original) and Parakeet (new) transcriptions.
View in HuggingFace Dataset Viewer
This dataset is viewable directly in the HuggingFace dataset viewer! Click the "Dataset Viewer" tab above to:
Listen to audio samples
See full Whisper and Parakeet transcriptions (not truncated)
Filter by CER bin… See the full description on the dataset page: https://huggingface.co/datasets/toth235a/voxpopuli-qc-samples-v3.Voxpopuli-ItVoxPopuli-Platinum-en-full
VoxPopuli-Platinum-en-full
Complete 142,066-row / ~404-hour English VoxPopuli Platinum dataset.
Reach out to data@trelis.com to purchase access or discuss a larger
custom-curation engagement.
Training Results
These results show why the Platinum labels matter. We compare the base model,
fine-tuning on raw VoxPopuli transcripts, and fine-tuning on this internally
filtered Platinum dataset. Evaluation uses the same english-spoken corpus WER
setup across four… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/VoxPopuli-Platinum-en-full.voxpopuli
Dataset Card for Voxpopuli
Dataset Summary
VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.
The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials.
This implementation contains transcribed speech data for 18 languages.
It also contains 29 hours of transcribed speech data of non-native English… See the full description on the dataset page: https://huggingface.co/datasets/Prabh139/voxpopuli.VoxPopuliAccentPairClassificationvoxpopuli_es-ja
Dataset Card for Spanish-to-Japanese Automatic Speech Recognition Dataset
Dataset Summary
This dataset was created as part of a workshop organized by Yasmin Moslem, focusing on speech-to-text pipelines.
The workshop's primary goal is to enable accurate transcription and translation of spoken a source language into a written language (and learn how to do so, of course 😃)
The dataset serves as the foundation for developing and evaluating various models, including… See the full description on the dataset page: https://huggingface.co/datasets/marianogonzalezgomez/voxpopuli_es-ja.voxpopuli-qc-samples-v2slue-voxpopuli
Dataset Card for "slue-voxpopuli"
More Information needed
voxpopuli_fi_processed_audio
Dataset Card for "voxpopuli_fi_processed_audio"
More Information needed
voxpopuli_fi_pseudo_labelledvoxpopuli-huvoxpopuli_fr_pseudo_labelledvoxpopuli_hu_dac_pairs
🎙️ VoxPopuli Hungarian DAC Speaker Consistency Dataset
Ez az adathalmaz a facebook/voxpopuli magyar nyelvű szeletéből készült, kifejezetten audio-nyelvi modellek (pl. DAC token generátorok) tanításához és finomhangolásához.
🎯 Koncepció: Speaker-Consistency
A dataset elsődleges célja a beszélő-konzisztens generálás tanítása. A felépítése biztosítja, hogy a modell megtanulja egy adott beszélő hangszínét (timbre) átvinni egy új szövegre.
Párosítási logika: A rendszer… See the full description on the dataset page: https://huggingface.co/datasets/fablevi/voxpopuli_hu_dac_pairs.voxpopuli-accent-clustering
