datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
german-commons
German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
A comprehensive collection of German-language text data under open licenses for training German language models.
Datasheet: DATASHEET.md.
Paper: arxiv.org/abs/2510.13996
Code: github.com/coral-nlp/llmdata
Bloom Filter (DOLMA-compatible): bloom_filter.bin
Dataset Description
This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.german-commons
German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
A comprehensive collection of German-language text data under open licenses for training German language models.
Datasheet: DATASHEET.md.
Paper: arxiv.org/abs/2510.13996
Code: github.com/coral-nlp/llmdata
Bloom Filter (DOLMA-compatible): bloom_filter.bin
Dataset Description
This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/CUI03/german-commons.German-PD
🇩🇪 German Public Domain 🇩🇪
German-Public Domain or German-PD is a large collection aiming to aggregate all German monographies and periodicals in the public domain. As of March 2024, it is the biggest German open corpus.
Dataset summary
The collection contains 260,638 individual texts making up 37,650,706,611 words recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/German-PD.CARD-Germany-Batch1
CARD – Germany 2 Days
A comprehensive multi-modal driving dataset with stereo cameras, LiDAR, and depth annotations.
Dataset Structure
This dataset contains 28 sequences across 1 region(s):
germany_2days: 28 sequences
Data Format
Each sequence contains:
img/: Stereo camera images (cam_0, cam_1)
raw/: Raw sensor data
labels/: YOLO-format annotations
export/: Trajectory and calibration data
agg_depth/: Aggregated depth point clouds… See the full description on the dataset page: https://huggingface.co/datasets/CARD-Data/CARD-Germany-Batch1.Aleph-Alpha-GermanWeb
AlephAlphaGermanWeb
Aleph-Alpha-GermanWeb is a new German-language dataset that combines heuristic and model-based filtering techniques with synthetic data generation to achieve SOTA performance in German-language benchmarks. The dataset draws from three sources: (1) Common Crawl web data, (2) FineWeb2, and (3) synthetically-generated data conditioned on actual, organic web data.
In our accompanying paper (published at EACL 2026), we evaluated our dataset by training both a 1B… See the full description on the dataset page: https://huggingface.co/datasets/Aleph-Alpha/Aleph-Alpha-GermanWeb.TransWeb-Edu-Germantranslated-german-english-asr
Translated German-English ASR Dataset
A large-scale, multi-source German speech dataset with paired English translations, designed for training and evaluating German Automatic Speech Recognition (ASR), Speech Translation, and Text-to-Speech (TTS) systems. This dataset is a curated mixture of well-established open-source German and multilingual speech corpora, all unified under a common schema with German audio, original German transcriptions, and English translations.… See the full description on the dataset page: https://huggingface.co/datasets/aman4014/translated-german-english-asr.german-canary-asr-0324
Dataset Beschreibung
Allgemeine Informationen
Dieser Datensatz ist eine Kombination aus drei verschiedenen Quellen für die deutsche Sprache: Commonvoice 16.1, Voxpopuli und Multilingual librispeech. Die Daten wurden gefiltert, normalisiert und grammatikalisch korrigiert.
Die drei Datensätze wurden erneut transkribiert und mit den entsprechenden Audio-Daten abgeglichen, um genaue Transkriptionen zu erhalten. Anschließend wurde ein Abgleich mit den Originaltranskripten… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/german-canary-asr-0324.laions_got_talent_german_bicodeccourt-decisions-germany
Open Legal Data: Court Decisions Germany
This dataset is a preprocessed version of an Open Legal Data data dump, spefically it contains German court decisions.
The dataset was automatically generated and uploaded to the HF hub using oldp-toolkit.
Available dumps
Date
Configs
2026-05-20
dump-20260520, dump-20260520-10k, dump-20260520-1k
2022-10-18
dump-20221018, dump-20221018-10k, dump-20221018-1k
Data format
Each dataset sample has the… See the full description on the dataset page: https://huggingface.co/datasets/openlegaldata/court-decisions-germany.ipfs_germany_laws_ir
Germany legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_germany_laws (revision 62477e216917df268f186972ef49581b02144156) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Germany prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_germany_laws_ir.german-asr-mixed-whisper
Dataset Card
Dataset Sources and Licensing
This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use.
Dataset Name
Original Source / Author
Link
TUDA-De German Speech Corpus
LT Group at UHH / TU Darmstadt
https://huggingface.co/datasets/uhhlt/Tuda-De
Mozilla Common Voice
Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/fosple/german-asr-mixed-whisper.TTS-German
TTS-German
High-quality German speech dataset for TTS and ASR, derived from CML-TTS German.
Processing Pipeline
Standardize → 24kHz mono WAV, loudness normalize
Transcribe → WhisperX word-level timestamps
Segment → ≤12s at word boundaries
Denoise → DeepFilterNet
Quality filter → DNSMOS ≥ 2.5
G2P → IPA phonemes (custom dictionary)
Statistics
Metric
Value
Samples
670,509
Hours
1250h
Sample rate
24kHz mono
Max duration
12s
Schema… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-German.german-wikipedia-articlesnanochat-german-data
nanochat German: Pretraining Data
This repository uses a subset of the LLäMmlein pretraining dataset, which itself is a strict subset of the German portion of the RedPajama V2 dataset.
To construct the dataset, we download the first 20 JSONL files from LLäMmlein and merge them into a single collection.
Following the original nanochat dataset construction process, we then split the data into shards containing approximately 250 million characters each.
GermanSTSBenchmark
GermanSTSBenchmark
An MTEB dataset
Massive Text Embedding Benchmark
Semantic Textual Similarity Benchmark (STSbenchmark) dataset translated into German. Translations were originally done by T-Systems on site services GmbH.
Task category
t2t
DomainsNone
Reference
https://github.com/t-systems-on-site-services-gmbh/german-STSbenchmark
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb… See the full description on the dataset page: https://huggingface.co/datasets/mteb/GermanSTSBenchmark.multilingual-librispeech-german-labeledmultitask_german_examples_32kfinqa-germanasr-german-mixed
Dataset Beschreibung
Allgemeine Informationen
Dieser Datensatz ist eine Kombination aus drei verschiedenen Quellen für die deutsche Sprache: Commonvoice 17.0 und Multilingual librispeech. Die Daten wurden gefiltert, normalisiert und grammatikalisch korrigiert.
Die drei Datensätze wurden erneut transkribiert und mit den entsprechenden Audio-Daten abgeglichen, um genaue Transkriptionen zu erhalten. Anschließend wurde ein Abgleich mit den Originaltranskripten durchgeführt… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/asr-german-mixed.Romansh_German_Parallel_Data
Romansh–German Parallel Dataset (FineWeb-Based)
This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction.
Description
This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.German-PD-Newspapers
Dataset Card for Public Domain Newspapers (German)
This dataset contains 13 billion words of OCR text extracted from German historical newspapers.
Dataset Details
Dataset Description
Curated by: Sebastian Majstorovic
Language(s) (NLP): German
License: Dataset: CC0, Texts: Public Domain
Dataset Sources [optional]
Repository: https://www.deutsche-digitale-bibliothek.de/newspaper
Copyright & License
The newspapers texts have been… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/German-PD-Newspapers.SLR-Bench-German
🧠 SLR-Bench-German: Scalable Logical Reasoning Benchmark (German Edition)
SLR-Bench Multilingual Versions:
SLR-Bench-German is the German-language pendant of the original SLR-Benchdataset.
It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into German.
This enables systematic evaluation and training of Large Language Models (LLMs) in logical… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-German.kangaroo_dataset
German Kangaroo Benchmark
The complete German Mathematical Kangaroo archive from 1998 to 2025 as a
multiple-choice benchmark: 3,886 items from 140 exams in five grade groups
(3--4, 5--6, 7--8, 9--10, 11--13), worth 3, 4, or 5 points each. 1,746 items are
multimodal, with a question diagram, image-based answer options, or both. The
accompanying paper describes the extraction, the evaluation protocol, and the
results.
Files
kangaroo.parquet: the benchmark, 3,886… See the full description on the dataset page: https://huggingface.co/datasets/kangaroo-dataset-german/kangaroo_dataset.german_mgermanquad-retrieval
GermanQuAD-Retrieval
An MTEB dataset
Massive Text Embedding Benchmark
Context Retrieval for German Question Answering
Task category
t2t
Domains
Written, Non-fiction, Web
Reference
https://huggingface.co/datasets/deepset/germanquad
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["GermanQuAD-Retrieval"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/germanquad-retrieval.german-sohee-synthetic-tts-24k
Flevi Restiti Vici
Bis uns der Weg nur noch nach vorn blieb
This repository contains a synthetic German audiobook and text-to-speech training dataset based on the original novel:
Flevi Restiti ViciBis uns der Weg nur noch nach vorn blieb
The novel was written in German by Maurice Hartmann, who is the author and copyright holder.
Work in progress
This dataset is a work in progress. Future revisions may include corrected transcripts, regenerated audio… See the full description on the dataset page: https://huggingface.co/datasets/Muckylixx/german-sohee-synthetic-tts-24k.CXM_Arena_German
Dataset Card for CXM Arena German Benchmark Suite
Dataset Description
This dataset, "CXM Arena German Benchmark Suite," is a comprehensive collection designed to evaluate various AI capabilities within the Customer Experience Management (CXM) domain, specifically for the German language. It is closely modeled after the original CXM_Arena benchmark, but all data is in German. The suite consolidates five distinct tasks into a unified benchmark, enabling robust testing of… See the full description on the dataset page: https://huggingface.co/datasets/sprinklr-huggingface/CXM_Arena_German.german-traffic-sign-detection
Dataset Labels
['animals', 'construction', 'cycles crossing', 'danger', 'no entry', 'pedestrian crossing', 'school crossing', 'snow', 'stop', 'bend', 'bend left', 'bend right', 'give way', 'go left', 'go left or straight', 'go right', 'go right or straight', 'go straight', 'keep left', 'keep right', 'no overtaking', 'no overtaking -trucks-', 'no traffic both ways', 'no trucks', 'priority at next intersection', 'priority road', 'restriction ends', 'restriction ends -overtaking… See the full description on the dataset page: https://huggingface.co/datasets/keremberke/german-traffic-sign-detection.german-courts
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/rusheeliyer/german-courts.
