CUI03/german-commons
German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokens of German text… See the full description on the dataset page: https://huggingface.co/datasets/CUI03/german-commons.
13k
Duplicate from coral-nlp/german-commons
