CoolFace
Datasetpublicgated

CLAUSE-Bielefeld/10M_phonemized_German_dataset

Small Phonemized German dataset Description This dataset is a smaller and phonemized version of Bastian Bunzeck's dataset (Bunzeck et al., 2025) and contains approximately 10 million words. The phonemization was done using Phonemizer by Bernard & Titeux, 2021. This dataset was used to train the Phoneme-based German babyLlama model. Source Composition The dataset has been compiled from the following sources: Source Description Words… See the full description on the dataset page: https://huggingface.co/datasets/CLAUSE-Bielefeld/10M_phonemized_German_dataset.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes7downloads

CLAUSE-Bielefeld/10M_phonemized_German_dataset · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.