CoolFace
Datasetpublic

CUI03/german-commons

German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokens of German text… See the full description on the dataset page: https://huggingface.co/datasets/CUI03/german-commons.

sourceHugging Faceodc-byupdated 9mo agoView on Hugging Face
1likes3kdownloads
1 commits on main
998461e9mo ago

Duplicate from coral-nlp/german-commons

CUI03, lgienapp