CoolFace
Datasetpublic

facebook/voxpopuli

Dataset Card for Voxpopuli Dataset Summary VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials. This implementation contains transcribed speech data for 18 languages. It also contains 29 hours of transcribed speech data of non-native… See the full description on the dataset page: https://huggingface.co/datasets/facebook/voxpopuli.

sourceHugging Facecc0-1.0updated 8mo agoView on Hugging Face
164likes76kdownloads

facebook/voxpopuli · main · files are served by the source, never re-hosted here