CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01facebook /voxpopuli Dataset Card for Voxpopuli Dataset Summary VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials. This implementation contains transcribed speech data for 18 languages. It also contains 29 hours of transcribed speech data of non-native… See the full description on the dataset page: https://huggingface.co/datasets/facebook/voxpopuli.audioautomatic-speech-recognition1M<n<10M164 likes76k downloads8mo agoHugging Face02ArtificialAnalysis /VoxPopuli-Cleaned-AA VoxPopuli-Cleaned-AA Quick links: AA Speech to Text Leaderboard | AA-WER v2.0 article VoxPopuli-Cleaned-AA is a cleaned subset of the English VoxPopuli test data from esb/datasets, a speech dataset derived from European Parliament recordings. This cleaned subset is the VoxPopuli portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation of Speech to Text (STT) models. This dataset is part of AA-WER… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/VoxPopuli-Cleaned-AA.audioautomatic-speech-recognitionn<1K7 likes1.2k downloads7mo agoHugging Face03kfajdsl /voxpopuliA large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.automatic-speech-recognition0 likes369 downloads9mo agoHugging Face04Prabh139 /voxpopuli Dataset Card for Voxpopuli Dataset Summary VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials. This implementation contains transcribed speech data for 18 languages. It also contains 29 hours of transcribed speech data of non-native English… See the full description on the dataset page: https://huggingface.co/datasets/Prabh139/voxpopuli.audioautomatic-speech-recognition1M<n<10M0 likes109 downloads5mo agoHugging Face05Trelis /VoxPopuli-Platinum-en-full VoxPopuli-Platinum-en-full Complete 142,066-row / ~404-hour English VoxPopuli Platinum dataset. Reach out to data@trelis.com to purchase access or discuss a larger custom-curation engagement. Training Results These results show why the Platinum labels matter. We compare the base model, fine-tuning on raw VoxPopuli transcripts, and fine-tuning on this internally filtered Platinum dataset. Evaluation uses the same english-spoken corpus WER setup across four… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/VoxPopuli-Platinum-en-full.audioautomatic-speech-recognition100K<n<1M0 likes104 downloads4mo agoHugging Face06marianogonzalezgomez /voxpopuli_es-ja Dataset Card for Spanish-to-Japanese Automatic Speech Recognition Dataset Dataset Summary This dataset was created as part of a workshop organized by Yasmin Moslem, focusing on speech-to-text pipelines. The workshop's primary goal is to enable accurate transcription and translation of spoken a source language into a written language (and learn how to do so, of course 😃) The dataset serves as the foundation for developing and evaluating various models, including… See the full description on the dataset page: https://huggingface.co/datasets/marianogonzalezgomez/voxpopuli_es-ja.audioautomatic-speech-recognition10K<n<100K1 likes62 downloads2y agoHugging Face07TheNHz /ellipsis-voxpopuli-ep-mouth-roigated Ellipsis — European Parliament Mouth-ROI (multilingual VSR) Mouth-region-of-interest (ROI) clips + verbatim transcripts for visual speech recognition (VSR / lip-reading), derived from European Parliament plenary debates (2019–2020) through the CC0 VoxPopuli segment index. 57,735 clips across 16 languages, in 58 WebDataset shards (~106 GB). Each clip is the speaker's mouth ROI only — a [T, 96, 96] grayscale tensor (Auto-AVSR convention) — paired with the VoxPopuli verbatim… See the full description on the dataset page: https://huggingface.co/datasets/TheNHz/ellipsis-voxpopuli-ep-mouth-roi.automatic-speech-recognition10K<n<100K0 likes33 downloads1mo agoHugging Face08distil-whisper /voxpopuliA large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.automatic-speech-recognition1 likes27 downloads3y agoHugging Face09Trelis /VoxPopuli-Platinum-en VoxPopuli-Platinum-en Showcase page and playable sample rows for the full 142,066-row / ~404-hour English VoxPopuli Platinum dataset. Reach out to data@trelis.com to purchase access or discuss a larger custom-curation engagement. Training Results These results show why the Platinum labels matter. We compare the base model, fine-tuning on raw VoxPopuli transcripts, and fine-tuning on this internally filtered Platinum dataset. Evaluation uses the same… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/VoxPopuli-Platinum-en.audioautomatic-speech-recognitionn<1K0 likes23 downloads4mo agoHugging Face10ggfox00000 /stt-voxpopuli-test-fr VoxPopuli — French test split (mirror of facebook/voxpopuli) Mirror public du split test de VoxPopuli config fr (Wang et al. ACL 2021, Facebook AI), pour benchmark ASR français parlementaire (sessions du Parlement européen, 2009-2020, locuteurs MEP variés). Ce repo ne contient que le split test FR (1 742 utterances ≈ 4-5 h). Pour les splits train / validation ou les autres langues, voir le repo upstream facebook/voxpopuli. Contenu 1 742 utterances FR (eurodéputés… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-voxpopuli-test-fr.audioautomatic-speech-recognition1K<n<10K0 likes20 downloads5mo agoHugging Face11Philou134 /voxpopuli_es-ja Dataset Card for Spanish-to-Japanese Automatic Speech Recognition Dataset Dataset Summary This dataset was created as part of a workshop organized by Yasmin Moslem, focusing on speech-to-text pipelines. The workshop's primary goal is to enable accurate transcription and translation of spoken a source language into a written language (and learn how to do so, of course 😃) The dataset serves as the foundation for developing and evaluating various models, including… See the full description on the dataset page: https://huggingface.co/datasets/Philou134/voxpopuli_es-ja.audioautomatic-speech-recognition10K<n<100K0 likes18 downloads2mo agoHugging Face12distil-whisper /voxpopuli-timestampedA large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.automatic-speech-recognition0 likes13 downloads3y agoHugging Face13ggfox00000 /stt-voxpopuli-test-en VoxPopuli EN — test split Mirror byte-exact du split test de facebook/voxpopuli config en. 1842 utt parlementaires (Parlement européen, 2009–2020) 16 kHz mono WAV embedded Référence WER : normalized_text (préférée à raw_text) Champ is_gold_transcript : True ⇒ transcription validée humaine Source : Wang, Riviere et al. VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation, ACL 2021. Usage from… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-voxpopuli-test-en.audioautomatic-speech-recognition1K<n<10K0 likes12 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.