CoolFace
Datasetpublic

MultiLlasa/Kartoffelphon-2.5M-de-ger

Kartoffelphon-2.5M-de-ger Kartoffelphon-2.5M-de-ger is a large-scale German speech dataset built as foundation data for Kartoffel TTS models and related Kartoffel speech projects. The dataset contains approximately 2.5 million audio-text snippets and an estimated 7,000 hours of speech. The data is mainly German, with some English segments intentionally retained. Dataset Summary The dataset is composed mostly of CC / CC-BY based podcast audio, with additional… See the full description on the dataset page: https://huggingface.co/datasets/MultiLlasa/Kartoffelphon-2.5M-de-ger.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
6likes964downloads
Dataset Card

Kartoffelphon-2.5M-de-ger

[image]

Kartoffelphon-2.5M-de-ger is a large-scale German speech dataset built as foundation data for Kartoffel TTS models and related Kartoffel speech projects.

The dataset contains approximately 2.5 million audio-text snippets and an estimated 7,000 hours of speech. The data is mainly German, with some English segments intentionally retained.

Dataset Summary

The dataset is composed mostly of CC / CC-BY based podcast audio, with additional material from sources such as LibriVox and lectures. LibriVox material was processed directly from source audio and was not taken from existing prepared speech datasets.

Audio was processed with a slightly adjusted Emilia-style pipeline, including audio cleaning, segmentation, and Whisper-based transcription.

Data Sources

The corpus is mainly based on:

  • —CC / CC-BY podcast audio
  • —LibriVox recordings processed from source audio
  • —lectures and other long-form spoken material

The dataset is not limited to studio-quality speech. It intentionally includes varied real-world speech conditions, speakers, recording setups, and speaking styles.

Language

The dataset is primarily German.

Some English speech is present and intentionally retained.

Processing Pipeline

The data was processed using a modified Emilia-style speech processing pipeline:

  1. 1.Source audio collection from permissively licensed sources
  2. 2.Audio cleaning and quality filtering
  3. 3.Segmentation into shorter snippets
  4. 4.Whisper-based transcription
  5. 5.Metadata generation
  6. 6.DNSMOS-style quality estimation

Licensing

The dataset is released under CC-BY 4.0.

The source material is mainly based on Creative Commons / CC-BY compatible audio, including podcasts, LibriVox recordings processed from source audio, and lecture-style material.

Citation

If you use this dataset, please cite it as:

bibtex
@misc{kartoffelphon_2_5m_de_ger,
  title        = {Kartoffelphon-2.5M-de-ger},
  author       = {MultiLlasa},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/MultiLlasa/Kartoffelphon-2.5M-de-ger}}
}