datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden
Dataset Card for "The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden"
Homepage: Spoken Wikipedia CorporaRepository: Bitbucket RepositoryPaper: Publication at LREC 2016Leaderboard: Interspeech 2018 PaperPoint of Contact: nats@nats.gitlab.io
Dataset Summary
The Spoken Wikipedia Corpora (SWC) is a collection of aligned spoken Wikipedia articles including articles in Dutch. It includes approximately 210 hours of audio, transcriptions, and metadata. The corpus is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Jaspernl/The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden.Dutch-Speech-Dataset
🎧 Dutch Speech Dataset
The Dutch Speech Dataset is a high-quality speech audio dataset designed to provide structured and diverse audio data for modern AI and machine learning applications. It includes 179 hours of audio data across 548 files, delivered in MP3 and WAV formats, with a total size of 190 MB. This well-organized audio dataset ensures balanced and representative voice data, with 51% female and 49% male speakers, and a wide age distribution from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Dutch-Speech-Dataset.YodaLingua-Dutch
YodaLingua-Dutch
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Dutch portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
50,839 audio–transcription pairs
Total duration
139 hours
Speakers
2,761 distinct speakers
Audio format
MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Dutch.
