CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01coral-nlp /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.tabulartext-generation10M<n<100M41 likes3.3k downloads8mo agoHugging Face02CUI03 /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/CUI03/german-commons.tabulartext-generation10M<n<100M1 likes3k downloads9mo agoHugging Face03PleIAs /German-PD 🇩🇪 German Public Domain 🇩🇪 German-Public Domain or German-PD is a large collection aiming to aggregate all German monographies and periodicals in the public domain. As of March 2024, it is the biggest German open corpus. Dataset summary The collection contains 260,638 individual texts making up 37,650,706,611 words recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/German-PD.text100K<n<1M15 likes2.8k downloads2y agoHugging Face04CARD-Data /CARD-Germany-Batch1gated CARD – Germany 2 Days A comprehensive multi-modal driving dataset with stereo cameras, LiDAR, and depth annotations. Dataset Structure This dataset contains 28 sequences across 1 region(s): germany_2days: 28 sequences Data Format Each sequence contains: img/: Stereo camera images (cam_0, cam_1) raw/: Raw sensor data labels/: YOLO-format annotations export/: Trajectory and calibration data agg_depth/: Aggregated depth point clouds… See the full description on the dataset page: https://huggingface.co/datasets/CARD-Data/CARD-Germany-Batch1.imagedepth-estimation10K<n<100K5 likes1.4k downloads2mo agoHugging Face05Aleph-Alpha /Aleph-Alpha-GermanWeb AlephAlphaGermanWeb Aleph-Alpha-GermanWeb is a new German-language dataset that combines heuristic and model-based filtering techniques with synthetic data generation to achieve SOTA performance in German-language benchmarks. The dataset draws from three sources: (1) Common Crawl web data, (2) FineWeb2, and (3) synthetically-generated data conditioned on actual, organic web data. In our accompanying paper (published at EACL 2026), we evaluated our dataset by training both a 1B… See the full description on the dataset page: https://huggingface.co/datasets/Aleph-Alpha/Aleph-Alpha-GermanWeb.text1B<n<10B24 likes1.2k downloads6mo agoHugging Face06britllm /TransWeb-Edu-Germantext10M<n<100M1 likes1.1k downloads2y agoHugging Face07aman4014 /translated-german-english-asr Translated German-English ASR Dataset A large-scale, multi-source German speech dataset with paired English translations, designed for training and evaluating German Automatic Speech Recognition (ASR), Speech Translation, and Text-to-Speech (TTS) systems. This dataset is a curated mixture of well-established open-source German and multilingual speech corpora, all unified under a common schema with German audio, original German transcriptions, and English translations.… See the full description on the dataset page: https://huggingface.co/datasets/aman4014/translated-german-english-asr.audioautomatic-speech-recognition1M<n<10M4 likes1.1k downloads5mo agoHugging Face08GermanEval /germeval_14The GermEval 2014 NER Shared Task builds on a new dataset with German Named Entity annotation with the following properties: - The data was sampled from German Wikipedia and News Corpora as a collection of citations. - The dataset covers over 31,000 sentences corresponding to over 590,000 tokens. - The NER annotation uses the NoSta-D guidelines, which extend the Tübingen Treebank guidelines, using four main NER categories with sub-structure, and annotating embeddings among NEs such as [ORG FC Kickers [LOC Darmstadt]].token-classification100K<n<1M4 likes1.1k downloads3y agoHugging Face09flozi00 /german-canary-asr-0324 Dataset Beschreibung Allgemeine Informationen Dieser Datensatz ist eine Kombination aus drei verschiedenen Quellen für die deutsche Sprache: Commonvoice 16.1, Voxpopuli und Multilingual librispeech. Die Daten wurden gefiltert, normalisiert und grammatikalisch korrigiert. Die drei Datensätze wurden erneut transkribiert und mit den entsprechenden Audio-Daten abgeglichen, um genaue Transkriptionen zu erhalten. Anschließend wurde ein Abgleich mit den Originaltranskripten… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/german-canary-asr-0324.audioautomatic-speech-recognition100K<n<1M8 likes1k downloads3y agoHugging Face10datadriven-company /TTS-German TTS-German High-quality German speech dataset for TTS and ASR, derived from CML-TTS German. Processing Pipeline Standardize → 24kHz mono WAV, loudness normalize Transcribe → WhisperX word-level timestamps Segment → ≤12s at word boundaries Denoise → DeepFilterNet Quality filter → DNSMOS ≥ 2.5 G2P → IPA phonemes (custom dictionary) Statistics Metric Value Samples 670,509 Hours 1250h Sample rate 24kHz mono Max duration 12s Schema… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-German.audiotext-to-speech1M<n<10M4 likes1k downloads6mo agoHugging Face11tudarmstadt-lt /germanerGermaNER is a freely available statistical German Named Entity Tagger based on conditional random fields(CRF). The tagger is trained and evaluated on the NoSta-D Named Entity dataset, which was used in the GermEval 2014 for named entity recognition. The tagger comes close to the performance of the best (proprietary) system in the competition with 77% F-measure (this is the latest result; the one reported in the paper is 76%) test set performance on the four standard NER classes (PERson, LOCation, ORGanisation and OTHer). We describe a range of features and their influence on German NER classification and provide a comparative evaluation and some analysis of the results. The software components, the training data and all data used for feature generation are distributed under permissive licenses, thus this tagger can be used in academic and commercial settings without restrictions or fees. The tagger is available as a command-line tool and as an Apache UIMA component.token-classification10K<n<100K4 likes1k downloads3y agoHugging Face12laion /laions_got_talent_german_bicodecaudio100K<n<1M0 likes975 downloads2y agoHugging Face13openlegaldata /court-decisions-germanygated Open Legal Data: Court Decisions Germany This dataset is a preprocessed version of an Open Legal Data data dump, spefically it contains German court decisions. The dataset was automatically generated and uploaded to the HF hub using oldp-toolkit. Available dumps Date Configs 2026-05-20 dump-20260520, dump-20260520-10k, dump-20260520-1k 2022-10-18 dump-20221018, dump-20221018-10k, dump-20221018-1k Data format Each dataset sample has the… See the full description on the dataset page: https://huggingface.co/datasets/openlegaldata/court-decisions-germany.texttext-generation100K<n<1M15 likes933 downloads4mo agoHugging Face14justicedao /ipfs_germany_laws_ir justicedao/ipfs_germany_laws_ir Germany laws IR release. tabular1M<n<10M0 likes886 downloads11d agoHugging Face15jonas-is-coding /german-wikipedia-articlestext1M<n<10M2 likes879 downloads2y agoHugging Face16fosple /german-asr-mixed-whisper Dataset Card Dataset Sources and Licensing This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use. Dataset Name Original Source / Author Link TUDA-De German Speech Corpus LT Group at UHH / TU Darmstadt https://huggingface.co/datasets/uhhlt/Tuda-De Mozilla Common Voice Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/fosple/german-asr-mixed-whisper.audioautomatic-speech-recognition1M<n<10M0 likes842 downloads6mo agoHugging Face17stefan-it /nanochat-german-data nanochat German: Pretraining Data This repository uses a subset of the LLäMmlein pretraining dataset, which itself is a strict subset of the German portion of the RedPajama V2 dataset. To construct the dataset, we download the first 20 JSONL files from LLäMmlein and merge them into a single collection. Following the original nanochat dataset construction process, we then split the data into shards containing approximately 250 million characters each. text10M<n<100M0 likes774 downloads11mo agoHugging Face18mteb /GermanSTSBenchmark GermanSTSBenchmark An MTEB dataset Massive Text Embedding Benchmark Semantic Textual Similarity Benchmark (STSbenchmark) dataset translated into German. Translations were originally done by T-Systems on site services GmbH. Task category t2t DomainsNone Reference https://github.com/t-systems-on-site-services-gmbh/german-STSbenchmark How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb… See the full description on the dataset page: https://huggingface.co/datasets/mteb/GermanSTSBenchmark.textsentence-similarity1K<n<10K0 likes724 downloads1y agoHugging Face19HumynLabs /German_Documents_Dataset_PDFdocumentn<1K0 likes709 downloads11mo agoHugging Face20flozi00 /multilingual-librispeech-german-labeledaudio100K<n<1M1 likes527 downloads2y agoHugging Face21seedboxai /multitask_german_examples_32ktabular100K<n<1M15 likes517 downloads3y agoHugging Face22deepset /germanquadIn order to raise the bar for non-English QA, we are releasing a high-quality, human-labeled German QA dataset consisting of 13 722 questions, incl. a three-way annotated test set. The creation of GermanQuAD is inspired by insights from existing datasets as well as our labeling experience from several industry projects. We combine the strengths of SQuAD, such as high out-of-domain performance, with self-sufficient questions that contain all relevant information for open-domain QA as in the NaturalQuestions dataset. Our training and test datasets do not overlap like other popular datasets and include complex questions that cannot be answered with a single entity or only a few words.question-answering44 likes504 downloads3y agoHugging Face23storytracer /German-PD-Newspapers Dataset Card for Public Domain Newspapers (German) This dataset contains 13 billion words of OCR text extracted from German historical newspapers. Dataset Details Dataset Description Curated by: Sebastian Majstorovic Language(s) (NLP): German License: Dataset: CC0, Texts: Public Domain Dataset Sources [optional] Repository: https://www.deutsche-digitale-bibliothek.de/newspaper Copyright & License The newspapers texts have been… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/German-PD-Newspapers.texttext-generation1M<n<10M5 likes490 downloads3y agoHugging Face24bowang0911 /finqa-germantabular1K<n<10K0 likes480 downloads3mo agoHugging Face25flozi00 /asr-german-mixed Dataset Beschreibung Allgemeine Informationen Dieser Datensatz ist eine Kombination aus drei verschiedenen Quellen für die deutsche Sprache: Commonvoice 17.0 und Multilingual librispeech. Die Daten wurden gefiltert, normalisiert und grammatikalisch korrigiert. Die drei Datensätze wurden erneut transkribiert und mit den entsprechenden Audio-Daten abgeglichen, um genaue Transkriptionen zu erhalten. Anschließend wurde ein Abgleich mit den Originaltranskripten durchgeführt… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/asr-german-mixed.audioautomatic-speech-recognition100K<n<1M9 likes478 downloads2y agoHugging Face26Sudehsna /Romansh_German_Parallel_Data Romansh–German Parallel Dataset (FineWeb-Based) This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction. Description This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.tabular10K<n<100K2 likes472 downloads1y agoHugging Face27AIML-TUDA /SLR-Bench-German 🧠 SLR-Bench-German: Scalable Logical Reasoning Benchmark (German Edition) SLR-Bench Multilingual Versions: SLR-Bench-German is the German-language pendant of the original SLR-Benchdataset. It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into German. This enables systematic evaluation and training of Large Language Models (LLMs) in logical… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-German.tabular10K<n<100K2 likes418 downloads4mo agoHugging Face28datastuff /german_maudio100K<n<1M0 likes415 downloads1y agoHugging Face29iamfebin /german-news-intelligence0 likes393 downloads11h agoHugging Face30mteb /germanquad-retrieval GermanQuAD-Retrieval An MTEB dataset Massive Text Embedding Benchmark Context Retrieval for German Question Answering Task category t2t Domains Written, Non-fiction, Web Reference https://huggingface.co/datasets/deepset/germanquad How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["GermanQuAD-Retrieval"]) evaluator = mteb.MTEB(task) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/germanquad-retrieval.texttext-retrieval1K<n<10K0 likes373 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.