datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
german-commons
German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
A comprehensive collection of German-language text data under open licenses for training German language models.
Datasheet: DATASHEET.md.
Paper: arxiv.org/abs/2510.13996
Code: github.com/coral-nlp/llmdata
Bloom Filter (DOLMA-compatible): bloom_filter.bin
Dataset Description
This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.german-commons
German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
A comprehensive collection of German-language text data under open licenses for training German language models.
Datasheet: DATASHEET.md.
Paper: arxiv.org/abs/2510.13996
Code: github.com/coral-nlp/llmdata
Bloom Filter (DOLMA-compatible): bloom_filter.bin
Dataset Description
This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/CUI03/german-commons.German-PD
🇩🇪 German Public Domain 🇩🇪
German-Public Domain or German-PD is a large collection aiming to aggregate all German monographies and periodicals in the public domain. As of March 2024, it is the biggest German open corpus.
Dataset summary
The collection contains 260,638 individual texts making up 37,650,706,611 words recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/German-PD.CARD-Germany-Batch1
CARD – Germany 2 Days
A comprehensive multi-modal driving dataset with stereo cameras, LiDAR, and depth annotations.
Dataset Structure
This dataset contains 28 sequences across 1 region(s):
germany_2days: 28 sequences
Data Format
Each sequence contains:
img/: Stereo camera images (cam_0, cam_1)
raw/: Raw sensor data
labels/: YOLO-format annotations
export/: Trajectory and calibration data
agg_depth/: Aggregated depth point clouds… See the full description on the dataset page: https://huggingface.co/datasets/CARD-Data/CARD-Germany-Batch1.Aleph-Alpha-GermanWeb
AlephAlphaGermanWeb
Aleph-Alpha-GermanWeb is a new German-language dataset that combines heuristic and model-based filtering techniques with synthetic data generation to achieve SOTA performance in German-language benchmarks. The dataset draws from three sources: (1) Common Crawl web data, (2) FineWeb2, and (3) synthetically-generated data conditioned on actual, organic web data.
In our accompanying paper (published at EACL 2026), we evaluated our dataset by training both a 1B… See the full description on the dataset page: https://huggingface.co/datasets/Aleph-Alpha/Aleph-Alpha-GermanWeb.TransWeb-Edu-Germantranslated-german-english-asr
Translated German-English ASR Dataset
A large-scale, multi-source German speech dataset with paired English translations, designed for training and evaluating German Automatic Speech Recognition (ASR), Speech Translation, and Text-to-Speech (TTS) systems. This dataset is a curated mixture of well-established open-source German and multilingual speech corpora, all unified under a common schema with German audio, original German transcriptions, and English translations.… See the full description on the dataset page: https://huggingface.co/datasets/aman4014/translated-german-english-asr.germeval_14The GermEval 2014 NER Shared Task builds on a new dataset with German Named Entity annotation with the following properties: - The data was sampled from German Wikipedia and News Corpora as a collection of citations. - The dataset covers over 31,000 sentences corresponding to over 590,000 tokens. - The NER annotation uses the NoSta-D guidelines, which extend the Tübingen Treebank guidelines, using four main NER categories with sub-structure, and annotating embeddings among NEs such as [ORG FC Kickers [LOC Darmstadt]].german-canary-asr-0324
Dataset Beschreibung
Allgemeine Informationen
Dieser Datensatz ist eine Kombination aus drei verschiedenen Quellen für die deutsche Sprache: Commonvoice 16.1, Voxpopuli und Multilingual librispeech. Die Daten wurden gefiltert, normalisiert und grammatikalisch korrigiert.
Die drei Datensätze wurden erneut transkribiert und mit den entsprechenden Audio-Daten abgeglichen, um genaue Transkriptionen zu erhalten. Anschließend wurde ein Abgleich mit den Originaltranskripten… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/german-canary-asr-0324.TTS-German
TTS-German
High-quality German speech dataset for TTS and ASR, derived from CML-TTS German.
Processing Pipeline
Standardize → 24kHz mono WAV, loudness normalize
Transcribe → WhisperX word-level timestamps
Segment → ≤12s at word boundaries
Denoise → DeepFilterNet
Quality filter → DNSMOS ≥ 2.5
G2P → IPA phonemes (custom dictionary)
Statistics
Metric
Value
Samples
670,509
Hours
1250h
Sample rate
24kHz mono
Max duration
12s
Schema… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-German.germanerGermaNER is a freely available statistical German Named Entity Tagger based on conditional random fields(CRF). The tagger is trained and evaluated on the NoSta-D Named Entity dataset, which was used in the GermEval 2014 for named entity recognition. The tagger comes close to the performance of the best (proprietary) system in the competition with 77% F-measure (this is the latest result; the one reported in the paper is 76%) test set performance on the four standard NER classes (PERson, LOCation, ORGanisation and OTHer).
We describe a range of features and their influence on German NER classification and provide a comparative evaluation and some analysis of the results. The software components, the training data and all data used for feature generation are distributed under permissive licenses, thus this tagger can be used in academic and commercial settings without restrictions or fees. The tagger is available as a command-line tool and as an Apache UIMA component.laions_got_talent_german_bicodeccourt-decisions-germany
Open Legal Data: Court Decisions Germany
This dataset is a preprocessed version of an Open Legal Data data dump, spefically it contains German court decisions.
The dataset was automatically generated and uploaded to the HF hub using oldp-toolkit.
Available dumps
Date
Configs
2026-05-20
dump-20260520, dump-20260520-10k, dump-20260520-1k
2022-10-18
dump-20221018, dump-20221018-10k, dump-20221018-1k
Data format
Each dataset sample has the… See the full description on the dataset page: https://huggingface.co/datasets/openlegaldata/court-decisions-germany.ipfs_germany_laws_ir
justicedao/ipfs_germany_laws_ir
Germany laws IR release.
german-wikipedia-articlesgerman-asr-mixed-whisper
Dataset Card
Dataset Sources and Licensing
This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use.
Dataset Name
Original Source / Author
Link
TUDA-De German Speech Corpus
LT Group at UHH / TU Darmstadt
https://huggingface.co/datasets/uhhlt/Tuda-De
Mozilla Common Voice
Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/fosple/german-asr-mixed-whisper.nanochat-german-data
nanochat German: Pretraining Data
This repository uses a subset of the LLäMmlein pretraining dataset, which itself is a strict subset of the German portion of the RedPajama V2 dataset.
To construct the dataset, we download the first 20 JSONL files from LLäMmlein and merge them into a single collection.
Following the original nanochat dataset construction process, we then split the data into shards containing approximately 250 million characters each.
GermanSTSBenchmark
GermanSTSBenchmark
An MTEB dataset
Massive Text Embedding Benchmark
Semantic Textual Similarity Benchmark (STSbenchmark) dataset translated into German. Translations were originally done by T-Systems on site services GmbH.
Task category
t2t
DomainsNone
Reference
https://github.com/t-systems-on-site-services-gmbh/german-STSbenchmark
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb… See the full description on the dataset page: https://huggingface.co/datasets/mteb/GermanSTSBenchmark.German_Documents_Dataset_PDFmultilingual-librispeech-german-labeledmultitask_german_examples_32kgermanquadIn order to raise the bar for non-English QA, we are releasing a high-quality, human-labeled German QA dataset consisting of 13 722 questions, incl. a three-way annotated test set.
The creation of GermanQuAD is inspired by insights from existing datasets as well as our labeling experience from several industry projects. We combine the strengths of SQuAD, such as high out-of-domain performance, with self-sufficient questions that contain all relevant information for open-domain QA as in the NaturalQuestions dataset. Our training and test datasets do not overlap like other popular datasets and include complex questions that cannot be answered with a single entity or only a few words.German-PD-Newspapers
Dataset Card for Public Domain Newspapers (German)
This dataset contains 13 billion words of OCR text extracted from German historical newspapers.
Dataset Details
Dataset Description
Curated by: Sebastian Majstorovic
Language(s) (NLP): German
License: Dataset: CC0, Texts: Public Domain
Dataset Sources [optional]
Repository: https://www.deutsche-digitale-bibliothek.de/newspaper
Copyright & License
The newspapers texts have been… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/German-PD-Newspapers.finqa-germanasr-german-mixed
Dataset Beschreibung
Allgemeine Informationen
Dieser Datensatz ist eine Kombination aus drei verschiedenen Quellen für die deutsche Sprache: Commonvoice 17.0 und Multilingual librispeech. Die Daten wurden gefiltert, normalisiert und grammatikalisch korrigiert.
Die drei Datensätze wurden erneut transkribiert und mit den entsprechenden Audio-Daten abgeglichen, um genaue Transkriptionen zu erhalten. Anschließend wurde ein Abgleich mit den Originaltranskripten durchgeführt… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/asr-german-mixed.Romansh_German_Parallel_Data
Romansh–German Parallel Dataset (FineWeb-Based)
This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction.
Description
This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.SLR-Bench-German
🧠 SLR-Bench-German: Scalable Logical Reasoning Benchmark (German Edition)
SLR-Bench Multilingual Versions:
SLR-Bench-German is the German-language pendant of the original SLR-Benchdataset.
It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into German.
This enables systematic evaluation and training of Large Language Models (LLMs) in logical… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-German.german_mgerman-news-intelligencegermanquad-retrieval
GermanQuAD-Retrieval
An MTEB dataset
Massive Text Embedding Benchmark
Context Retrieval for German Question Answering
Task category
t2t
Domains
Written, Non-fiction, Web
Reference
https://huggingface.co/datasets/deepset/germanquad
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["GermanQuAD-Retrieval"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/germanquad-retrieval.
