CLAUSE-Bielefeld/10M_phonemized_German_dataset
Small Phonemized German dataset Description This dataset is a smaller and phonemized version of Bastian Bunzeck's dataset (Bunzeck et al., 2025) and contains approximately 10 million words. The phonemization was done using Phonemizer by Bernard & Titeux, 2021. This dataset was used to train the Phoneme-based German babyLlama model. Source Composition The dataset has been compiled from the following sources: Source Description Words… See the full description on the dataset page: https://huggingface.co/datasets/CLAUSE-Bielefeld/10M_phonemized_German_dataset.
This repository belongs to CLAUSE-Bielefeld on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
