CoolFace
Datasetpublic

datadriven-company/TTS-German

TTS-German High-quality German speech dataset for TTS and ASR, derived from CML-TTS German. Processing Pipeline Standardize → 24kHz mono WAV, loudness normalize Transcribe → WhisperX word-level timestamps Segment → ≤12s at word boundaries Denoise → DeepFilterNet Quality filter → DNSMOS ≥ 2.5 G2P → IPA phonemes (custom dictionary) Statistics Metric Value Samples 670,509 Hours 1250h Sample rate 24kHz mono Max duration 12s… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-German.

sourceHugging Facecc-by-4.0updated 7mo agoView on Hugging Face
4likes525downloads

datadriven-company/TTS-German · main · files are served by the source, never re-hosted here