Thorsten-Voice/TV-24kHz-2025.12-Neutral-FT-Mini-tokenised
Thorsten-Voice TV-24kHz-2025.12-Neutral-FT-Mini-tokenised Overview This dataset is the tokenised version of the TV-24kHz-2025.12-Neutral-FT-Mini dataset, prepared specifically for direct fine-tuning of Orpheus TTS Thorsten-Voice. It contains 60 tokenised German speech samples, optimised for fast experimentation and precise speaker adaptation. This dataset is intended for: Lightweight Orpheus TTS fine-tuning Speaker identity refinement Prosody and articulation… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-2025.12-Neutral-FT-Mini-tokenised.
license: cc0-1.0 language:
- de task_categories:
- text-to-speech tags:
- text-to-speech
- tts
- german
- voice
- fine-tuning
- tokenized
- orpheus-tts
- thorsten-voice ---
Thorsten-Voice TV-24kHz-2025.12-Neutral-FT-Mini-tokenised
Overview
This dataset is the tokenised version of the TV-24kHz-2025.12-Neutral-FT-Mini dataset, prepared specifically for direct fine-tuning of Orpheus TTS Thorsten-Voice.
It contains 60 tokenised German speech samples, optimised for fast experimentation and precise speaker adaptation.
This dataset is intended for:
- Lightweight Orpheus TTS fine-tuning
- Speaker identity refinement
- Prosody and articulation adjustments
Dataset Origin
- Source dataset: Thorsten-Voice/TV-24kHz-2025.12-Neutral-FT-Mini
- Speaker: Thorsten
- Language: German (de)
- Recording date: December 2025
Processing Details
- Audio resampled to 24,000 Hz
- Loudness normalized to -24dB
- Tokenised using Orpheus TTS preprocessing
- Ready-to-train token format
Tokenisation was performed using the Orpheus TTS codebase as of December 2025.
License
This dataset is released under the Creative Commons Zero (CC0 1.0) license.
Free for any use, including commercial and derivative works, without attribution requirements.
Related Projects
- Thorsten-Voice Project: https://www.Thorsten-Voice.de
- Orpheus TTS (GitHub): https://github.com/canopylabs/orpheus-tts
Notes
This dataset is designed for precise and controlled fine-tuning of my German Thorsten-Voice Orpheus TTS voices and complements the larger Thorsten-Voice training datasets.
