CoolFace
Datasetpublicgated

CLAUSE-Bielefeld/10M_phonemized_German_dataset

Small Phonemized German dataset Description This dataset is a smaller and phonemized version of Bastian Bunzeck's dataset (Bunzeck et al., 2025) and contains approximately 10 million words. The phonemization was done using Phonemizer by Bernard & Titeux, 2021. This dataset was used to train the Phoneme-based German babyLlama model. Source Composition The dataset has been compiled from the following sources: Source Description Words… See the full description on the dataset page: https://huggingface.co/datasets/CLAUSE-Bielefeld/10M_phonemized_German_dataset.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes7downloads
settings

This repository belongs to CLAUSE-Bielefeld on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

name10M_phonemized_German_dataset
visibilitypublic
licencenot set
gatedyes
ownerCLAUSE-Bielefeld
Account settings
CLAUSE-Bielefeld/10M_phonemized_German_dataset · CoolFace