CoolFace
Datasetpublic

CUI03/german-commons

German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokens of German text… See the full description on the dataset page: https://huggingface.co/datasets/CUI03/german-commons.

sourceHugging Faceodc-byupdated 9mo agoView on Hugging Face
1likes3kdownloads
Dataset Card

German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models

A comprehensive collection of German-language text data under open licenses for training German language models.

Dataset Description

This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokens of German text data with 35.78 million documents spanning 7 thematic domains:

  • 🌐 Web Commons: 19.89B tokens source from Wiki projects, online discussions, code repositories, social media posts, YouTube transcripts
  • 💬 Political Commons: 3.57B tokens sourced from parliamentary documents, speeches, protocols, political vocabulary
  • ⚖️ Legal Commons: 2.99B tokens sourced from court decisions, federal law, legal databases, EU legal documents
  • 📰 News Commons: 72.67B tokens sourced from historical and current newspapers archives
  • 🏦 Economics Commons: 0.11B tokens sourced from EU public tenders
  • 📚 Cultural Commons: 54.49B tokens sourced from cultural heritage collections
  • 🔬 Scientific Commons: 0.84B tokens sourced from scholarly papers, books, and technical journals

Dataset Features

Each record contains the following fields:

  • id: Unique identifier string, as per each documents' source dataset
  • source: Source dataset name
  • subset: Thematic subset (Cultural, Legal, Political, Scientific, News, Web, Economic)
  • text: Main text content; deduplicated, quality filtered, with consistent formatting and encoding. Can be split at newlines to obtain paragraph text.
  • license: List of applicable licenses for each document, given as canonical SPDX license URL.
  • num_tokens: GPT-2 token count
  • perplexity: Text perplexity measured with a KenLM model trained on German Wikipedia text
  • ocr_score: OCR quality score measured using OCRoscope

Dataset Usage

  • Load the entire dataset
python
  from datasets import load_dataset
  
  ds = load_dataset("coral-nlp/german-commons")
  • Load a thematic subset
python
  ds = load_dataset("coral-nlp/german-commons", "cultural")
  • Load individual source datasets
python
  wikipedia = load_dataset("coral-nlp/german-commons", "web", split="wikipedia")

Supported splits and constituent datasets are:

SubsetSplit KeyDataset NameDocsTokensLicenseText TypeSource
webwikipediaWikipedia2,930,2242,948,751,608CC-BY-SA-4.0Various🔗
webwikidiscussionsWikipedia Discussions8,349,0761,218,210,917CC-BY-SA-4.0Online Discussions🔗
webyoutubecommonsYouTube Commons2,809,71414,478,850,964VariousVideo Subtitles🔗
webonemillionpostsOne Million Posts Corpus946,08294,872,633CC-BY-4.0Online Discussions🔗
webthestackThe Stack (Markdown and TXT Subsets)421,4661,105,173,228VariousVarious🔗
politicalreichtagsprotokolleReichtagsprotokolle522703,495,637CC-BY-SA-4.0Parliamentary Protocols🔗
politicalgermanpoliticalspeechesGerman Political Speeches6,67829,409,655CC-BY-4.0Speech Transcripts🔗
politicalbtdrucksachenCorpus der Drucksachen des Deutschen Bundestages3,017528,769,669CC0-1.0Parliamentary Publications🔗
politicalbtplenarprotokolleCorpus der Plenarprotokolle des Deutschen Bundestages1,833316,034,708CC0-1.0Parliamentary Protocols🔗
politicaleurovocEuroVoc245,8381,988,111,462EUPLParliamentary Publications🔗
legalbundesrechtCorpus des Deutschen Bundesrechts3,2171,004,294CC0-1.0German Federal Laws🔗
legalopenlegaldataOpenLegalData249,9091,915,956,613CC0-1.0Court Decisions🔗
legalbfhCorpus der Entscheidungen des BFH10,88567,791,931CC0-1.0Court Decisions🔗
legalbgh20Entscheidungen des BGH in Strafsachen des 20. Jhd.36,06292,873,390CC0-1.0Court Decisions🔗
legalbghCorpus der Entscheidungen des BGH77,258292,832,709CC0-1.0Court Decisions🔗
legalbverfgCorpus der Entscheidungen des BVerfG8,02839,503,223CC0-1.0Court Decisions🔗
legalbpatgCorpus der Entscheidungen des BpatG30,705185,099,188CC0-1.0Court Decisions🔗
legalbverwgCorpus der Entscheidungen des BVerwG27,185123,487,739CC0-1.0Court Decisions🔗
legalbverfgaesCorpus der amtl. Entscheidungssammlung des BVerfG91924,427,294CC0-1.0Court Decisions🔗
legalbagCorpus der Entscheidungen des BAG5,62448,248,111CC0-1.0Court Decisions🔗
legaleurlexEurLEX64,934201,263,562CC-BY-4.0European Union Laws🔗
newszeitungsportalDeutsches Zeitungsportal8,076,16443,871,094,547CC0-1.0News Articles🔗
newseuropeanaEuropeana Newspapers3,256,34120,684,418,365CC0-1.0News Articles🔗
newsannoANNO1,910,2818,103,825,248CC0-1.0News Articles🔗
newswikinewsWikiNews23,26614,222,520CC-BY-4.0News Articles🔗
economictedeutendersTEDEUTenders57,214110,611,112CC0-1.0Procurement Notices🔗
culturaldibilitDiBiLit-Korpus2,062216,391,448CC-BY-SA-4.0Literature🔗
culturaldibiphilDiBiPhil-Korpus26932,151,997CC-BY-SA-4.0Literature🔗
culturalwikisourceWikisource240,689347,770,430CC-BY-SA-4.0Various🔗
culturalwikivoyageWikivoyage20,37042,025,478CC-BY-SA-4.0Travel🔗
culturalgermanpdGerman-PD123,59249,333,198,231CC0-1.0Literature🔗
culturalblbooksBLBooks3,7141,012,047,216CC0-1.0Literature🔗
culturalmoselMOSEL3,127,2033,181,917,752CC-BY-4.0Speech Transcripts🔗
culturalsbbfulltextsSBB Fulltexts2,605,569358,514,283CC-BY-4.0Literature🔗
culturalwikiquoteWikiquote8,6126,688,458CC-BY-4.0Quotes & Proverbs🔗
scientificwikibooksWikibooks346180,257,799CC-BY-SA-4.0Educational Books🔗
scientificpolyjournalDigitalisierung des Polytechnischen Journals27,29250,434,996CC-BY-SA-4.0Scholarly Papers🔗
scientificdoabDirectory of Open Access Books1,939166,920,321VariousScholarly Books🔗
scientificarxivarXiv8103,478VariousScholarly Papers🔗
scientificopenalexOpenAlex47,733413,632,648VariousEducational Content🔗
scientificwikiversityWikiversity16,37127,802,099CC-BY-SA-4.0Scholarly Papers🔗

Citation

If you use this dataset, please cite the correponding paper:

bibtex
@article{gienapp:2025d,
	title        = {{The German Commons -- 154 Billion Tokens of Openly Licensed Text for German Language Models}},
	author       = {Lukas Gienapp and
                    Christopher Schr\"oder and
                    Stefan Schweter and
                    Christopher Akiki and
                    Ferdinand Schlatt and
                    Arden Zimmermann and
                    Phillipe Gen\^et and
                    Martin Potthast},
	year         = 2025,
	month        = oct,
	journal      = {CoRR},
	volume       = {abs/2510.13996},
	url          = {https://arxiv.org/abs/2510.13996}
}

License

This dataset aggregation and metadata is released under ODC-BY license. Individual documents have their own specific licenses - please check the license field for each record.