datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TV-44kHz-Full
The "Thorsten-Voice" dataset
This truly open source (CC0 license) german (🇩🇪) voice dataset contains about 40 hours of transcribed voice recordings by Thorsten Müller,
a single male, native speaker in over 38.000 wave files.
Mono
Samplerate: 44.100Hz
Trimmed silence at begin/end
Denoised
Normalized to -24dB
Disclaimer
"Please keep in mind, I am not a professional speaker, just an open source speech technology enthusiast who donates his voice. I contribute my personal… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-44kHz-Full.TV-24kHz-2025.12-Neutral-FT-Mini
Thorsten-Voice TV-24kHz-2025.12-Neutral-FT-Mini
Overview
This dataset is a small, high-quality fine-tuning dataset created specifically for speaker refinement and voice matching in Orpheus TTS models.
It consists of 60 newly recorded German speech samples, spoken in a neutral, relaxed, everyday style, closely reflecting the natural speaking voice of the original speaker.
This dataset is intended for:
Speaker adaptation and voice refinement
Fine-tuning Orpheus TTS models… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-2025.12-Neutral-FT-Mini.TV-24kHz-Neutral
Thorsten-Voice TV-24kHz-Neutral Dataset
This dataset is a resampled version of the "TV-2022.10-Neutral" configuration from the original Thorsten-Voice TV-44kHz-Full dataset, converted from 44.1kHz to 24kHz sampling rate.
Dataset Description
The Thorsten-Voice dataset contains German speech recordings by Thorsten Müller, suitable for text-to-speech (TTS) training and other speech synthesis tasks.
Changes from Original
Sample Rate: Converted from… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-Neutral.TV-24kHz-Neutral-tokenised
Thorsten-Voice TV-24kHz-Neutral-tokenised
Overview
This dataset is a tokenised German text-to-speech dataset created for training and fine-tuning the Orpheus TTS model family.
It is based on approximately 12,000 speech recordings from the original Thorsten-Voice Dataset (2022.10) and has been resampled to 24 kHz and tokenised using Orpheus TTS preprocessing.
This dataset is intended for:
Training and fine-tuning Orpheus-based German TTS models
Research on neural speech… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-Neutral-tokenised.TV-24kHz-2025.12-Neutral-FT-Mini-tokenised
Thorsten-Voice TV-24kHz-2025.12-Neutral-FT-Mini-tokenised
Overview
This dataset is the tokenised version of the TV-24kHz-2025.12-Neutral-FT-Mini dataset, prepared specifically for direct fine-tuning of Orpheus TTS Thorsten-Voice.
It contains 60 tokenised German speech samples, optimised for fast experimentation and precise speaker adaptation.
This dataset is intended for:
Lightweight Orpheus TTS fine-tuning
Speaker identity refinement
Prosody and articulation adjustments… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-2025.12-Neutral-FT-Mini-tokenised.
