CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AITRADER /dutch-tts-labeled-complete Dutch TTS Dataset - Complete Labeled A comprehensive Dutch text-to-speech dataset with 596,508 audio samples totaling 234GB of audio data. Quick Preview The default config shows a 100-row sample for the dataset viewer. To access the full dataset, use the full config. Dataset Description This dataset contains Dutch speech recordings with rich metadata including: Emotion labels (neutral, happy, sad, angry) Speaker IDs (239,388 unique speakers)… See the full description on the dataset page: https://huggingface.co/datasets/AITRADER/dutch-tts-labeled-complete.audiotext-to-speech100K<n<1M0 likes713 downloads9mo agoHugging Face02AdoCleanCode /mls_dutchaudio100K<n<1M1 likes599 downloads7mo agoHugging Face03freds0 /cml_tts_dataset_dutchaudio100K<n<1M1 likes188 downloads2y agoHugging Face04Jaspernl /The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden Dataset Card for "The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden" Homepage: Spoken Wikipedia CorporaRepository: Bitbucket RepositoryPaper: Publication at LREC 2016Leaderboard: Interspeech 2018 PaperPoint of Contact: nats@nats.gitlab.io Dataset Summary The Spoken Wikipedia Corpora (SWC) is a collection of aligned spoken Wikipedia articles including articles in Dutch. It includes approximately 210 hours of audio, transcriptions, and metadata. The corpus is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Jaspernl/The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden.audioautomatic-speech-recognition10K<n<100K1 likes130 downloads2y agoHugging Face05Speech-data /Dutch-Speech-Dataset 🎧 Dutch Speech Dataset The Dutch Speech Dataset is a high-quality speech audio dataset designed to provide structured and diverse audio data for modern AI and machine learning applications. It includes 179 hours of audio data across 548 files, delivered in MP3 and WAV formats, with a total size of 190 MB. This well-organized audio dataset ensures balanced and representative voice data, with 51% female and 49% male speakers, and a wide age distribution from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Dutch-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes29 downloads6mo agoHugging Face06datadriven-company /TTS-Dutchaudio100K<n<1M1 likes28 downloads7mo agoHugging Face07Tundragoon /jeroen-dutch1666 audio1K<n<10K0 likes9 downloads8mo agoHugging Face08Thomcles /YodaLingua-Dutchgated YodaLingua-Dutch YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Dutch portion of the multilingual YodaLingua collection. 🧾 Dataset Overview Property Value Total clips 50,839 audio–transcription pairs Total duration 139 hours Speakers 2,761 distinct speakers Audio format MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Dutch.audiotext-to-speech10K<n<100K0 likes9 downloads5mo agoHugging Face09Tundragoon /lotte-dutch1724 audio1K<n<10K0 likes7 downloads8mo agoHugging Face10mariatepei /natural_accented_Dutchaudion<1K0 likes5 downloads2y agoHugging Face11shunyalabs /dutch-speech-datasetaudio1K<n<10K0 likes5 downloads1y agoHugging Face12NescIO1 /cml_dutch_24k_subset_largeraudio10K<n<100K0 likes5 downloads10mo agoHugging Face13AdoCleanCode /voxopopuli_dutchaudio10K<n<100K0 likes5 downloads7mo agoHugging Face14Juvoly /dutch-medical-setgated Dataset Card for "dutch-medical-set" More Information needed audio1K<n<10K0 likes3 downloads3y agoHugging Face15mariatepei /synthetic_accented_Dutchaudio1K<n<10K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.