datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden
Dataset Card for "The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden"
Homepage: Spoken Wikipedia CorporaRepository: Bitbucket RepositoryPaper: Publication at LREC 2016Leaderboard: Interspeech 2018 PaperPoint of Contact: nats@nats.gitlab.io
Dataset Summary
The Spoken Wikipedia Corpora (SWC) is a collection of aligned spoken Wikipedia articles including articles in Dutch. It includes approximately 210 hours of audio, transcriptions, and metadata. The corpus is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Jaspernl/The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden.wikipedia_spanish
Dataset Card for wikipedia_spanish
Dataset Summary
According to the project page of the WikiProject Spoken Wikipedia:
The WikiProject Spoken Wikipedia aims to produce recordings of Wikipedia articles being read aloud. Therefore, the WIKIPEDIA SPANISH CORPUS is a dataset created from the Spanish version of the WikiProject Spoken Wikipedia, called Wikipedia Grabada
The WIKIPEDIA SPANISH CORPUS aims to be used in the Automatic Speech Recognition (ASR) task. It is a gender… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/wikipedia_spanish.
