CoolFace
Datasetpublicgated

CLAUSE-Bielefeld/10M_phonemized_German_dataset

Small Phonemized German dataset Description This dataset is a smaller and phonemized version of Bastian Bunzeck's dataset (Bunzeck et al., 2025) and contains approximately 10 million words. The phonemization was done using Phonemizer by Bernard & Titeux, 2021. This dataset was used to train the Phoneme-based German babyLlama model. Source Composition The dataset has been compiled from the following sources: Source Description Words… See the full description on the dataset page: https://huggingface.co/datasets/CLAUSE-Bielefeld/10M_phonemized_German_dataset.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes7downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.