CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01coral-nlp /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.tabulartext-generation10M<n<100M41 likes3.3k downloads8mo agoHugging Face02CUI03 /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/CUI03/german-commons.tabulartext-generation10M<n<100M1 likes3k downloads9mo agoHugging Face03PleIAs /German-PD 🇩🇪 German Public Domain 🇩🇪 German-Public Domain or German-PD is a large collection aiming to aggregate all German monographies and periodicals in the public domain. As of March 2024, it is the biggest German open corpus. Dataset summary The collection contains 260,638 individual texts making up 37,650,706,611 words recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/German-PD.text100K<n<1M15 likes2.9k downloads2y agoHugging Face04CARD-Data /CARD-Germany-Batch1gated CARD – Germany 2 Days A comprehensive multi-modal driving dataset with stereo cameras, LiDAR, and depth annotations. Dataset Structure This dataset contains 28 sequences across 1 region(s): germany_2days: 28 sequences Data Format Each sequence contains: img/: Stereo camera images (cam_0, cam_1) raw/: Raw sensor data labels/: YOLO-format annotations export/: Trajectory and calibration data agg_depth/: Aggregated depth point clouds… See the full description on the dataset page: https://huggingface.co/datasets/CARD-Data/CARD-Germany-Batch1.imagedepth-estimation10K<n<100K5 likes1.3k downloads2mo agoHugging Face05Aleph-Alpha /Aleph-Alpha-GermanWeb AlephAlphaGermanWeb Aleph-Alpha-GermanWeb is a new German-language dataset that combines heuristic and model-based filtering techniques with synthetic data generation to achieve SOTA performance in German-language benchmarks. The dataset draws from three sources: (1) Common Crawl web data, (2) FineWeb2, and (3) synthetically-generated data conditioned on actual, organic web data. In our accompanying paper (published at EACL 2026), we evaluated our dataset by training both a 1B… See the full description on the dataset page: https://huggingface.co/datasets/Aleph-Alpha/Aleph-Alpha-GermanWeb.text1B<n<10B24 likes1.2k downloads6mo agoHugging Face06britllm /TransWeb-Edu-Germantext10M<n<100M1 likes1.1k downloads2y agoHugging Face07aman4014 /translated-german-english-asr Translated German-English ASR Dataset A large-scale, multi-source German speech dataset with paired English translations, designed for training and evaluating German Automatic Speech Recognition (ASR), Speech Translation, and Text-to-Speech (TTS) systems. This dataset is a curated mixture of well-established open-source German and multilingual speech corpora, all unified under a common schema with German audio, original German transcriptions, and English translations.… See the full description on the dataset page: https://huggingface.co/datasets/aman4014/translated-german-english-asr.audioautomatic-speech-recognition1M<n<10M4 likes1.1k downloads5mo agoHugging Face08flozi00 /german-canary-asr-0324 Dataset Beschreibung Allgemeine Informationen Dieser Datensatz ist eine Kombination aus drei verschiedenen Quellen für die deutsche Sprache: Commonvoice 16.1, Voxpopuli und Multilingual librispeech. Die Daten wurden gefiltert, normalisiert und grammatikalisch korrigiert. Die drei Datensätze wurden erneut transkribiert und mit den entsprechenden Audio-Daten abgeglichen, um genaue Transkriptionen zu erhalten. Anschließend wurde ein Abgleich mit den Originaltranskripten… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/german-canary-asr-0324.audioautomatic-speech-recognition100K<n<1M8 likes1k downloads3y agoHugging Face09laion /laions_got_talent_german_bicodecaudio100K<n<1M0 likes977 downloads2y agoHugging Face10openlegaldata /court-decisions-germanygated Open Legal Data: Court Decisions Germany This dataset is a preprocessed version of an Open Legal Data data dump, spefically it contains German court decisions. The dataset was automatically generated and uploaded to the HF hub using oldp-toolkit. Available dumps Date Configs 2026-05-20 dump-20260520, dump-20260520-10k, dump-20260520-1k 2022-10-18 dump-20221018, dump-20221018-10k, dump-20221018-1k Data format Each dataset sample has the… See the full description on the dataset page: https://huggingface.co/datasets/openlegaldata/court-decisions-germany.texttext-generation100K<n<1M15 likes946 downloads4mo agoHugging Face11justicedao /ipfs_germany_laws_ir Germany legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_germany_laws (revision 62477e216917df268f186972ef49581b02144156) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Germany prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_germany_laws_ir.tabulartext-retrieval1M<n<10M0 likes888 downloads14h agoHugging Face12fosple /german-asr-mixed-whisper Dataset Card Dataset Sources and Licensing This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use. Dataset Name Original Source / Author Link TUDA-De German Speech Corpus LT Group at UHH / TU Darmstadt https://huggingface.co/datasets/uhhlt/Tuda-De Mozilla Common Voice Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/fosple/german-asr-mixed-whisper.audioautomatic-speech-recognition1M<n<10M0 likes867 downloads7mo agoHugging Face13datadriven-company /TTS-German TTS-German High-quality German speech dataset for TTS and ASR, derived from CML-TTS German. Processing Pipeline Standardize → 24kHz mono WAV, loudness normalize Transcribe → WhisperX word-level timestamps Segment → ≤12s at word boundaries Denoise → DeepFilterNet Quality filter → DNSMOS ≥ 2.5 G2P → IPA phonemes (custom dictionary) Statistics Metric Value Samples 670,509 Hours 1250h Sample rate 24kHz mono Max duration 12s Schema… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-German.audiotext-to-speech1M<n<10M4 likes815 downloads6mo agoHugging Face14jonas-is-coding /german-wikipedia-articlestext1M<n<10M2 likes799 downloads2y agoHugging Face15stefan-it /nanochat-german-data nanochat German: Pretraining Data This repository uses a subset of the LLäMmlein pretraining dataset, which itself is a strict subset of the German portion of the RedPajama V2 dataset. To construct the dataset, we download the first 20 JSONL files from LLäMmlein and merge them into a single collection. Following the original nanochat dataset construction process, we then split the data into shards containing approximately 250 million characters each. text10M<n<100M0 likes754 downloads11mo agoHugging Face16mteb /GermanSTSBenchmark GermanSTSBenchmark An MTEB dataset Massive Text Embedding Benchmark Semantic Textual Similarity Benchmark (STSbenchmark) dataset translated into German. Translations were originally done by T-Systems on site services GmbH. Task category t2t DomainsNone Reference https://github.com/t-systems-on-site-services-gmbh/german-STSbenchmark How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb… See the full description on the dataset page: https://huggingface.co/datasets/mteb/GermanSTSBenchmark.textsentence-similarity1K<n<10K0 likes735 downloads1y agoHugging Face17flozi00 /multilingual-librispeech-german-labeledaudio100K<n<1M1 likes527 downloads2y agoHugging Face18seedboxai /multitask_german_examples_32ktabular100K<n<1M15 likes524 downloads3y agoHugging Face19bowang0911 /finqa-germantabular1K<n<10K0 likes489 downloads3mo agoHugging Face20flozi00 /asr-german-mixed Dataset Beschreibung Allgemeine Informationen Dieser Datensatz ist eine Kombination aus drei verschiedenen Quellen für die deutsche Sprache: Commonvoice 17.0 und Multilingual librispeech. Die Daten wurden gefiltert, normalisiert und grammatikalisch korrigiert. Die drei Datensätze wurden erneut transkribiert und mit den entsprechenden Audio-Daten abgeglichen, um genaue Transkriptionen zu erhalten. Anschließend wurde ein Abgleich mit den Originaltranskripten durchgeführt… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/asr-german-mixed.audioautomatic-speech-recognition100K<n<1M9 likes478 downloads2y agoHugging Face21Sudehsna /Romansh_German_Parallel_Data Romansh–German Parallel Dataset (FineWeb-Based) This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction. Description This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.tabular10K<n<100K2 likes469 downloads1y agoHugging Face22storytracer /German-PD-Newspapers Dataset Card for Public Domain Newspapers (German) This dataset contains 13 billion words of OCR text extracted from German historical newspapers. Dataset Details Dataset Description Curated by: Sebastian Majstorovic Language(s) (NLP): German License: Dataset: CC0, Texts: Public Domain Dataset Sources [optional] Repository: https://www.deutsche-digitale-bibliothek.de/newspaper Copyright & License The newspapers texts have been… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/German-PD-Newspapers.texttext-generation1M<n<10M5 likes428 downloads3y agoHugging Face23AIML-TUDA /SLR-Bench-German 🧠 SLR-Bench-German: Scalable Logical Reasoning Benchmark (German Edition) SLR-Bench Multilingual Versions: SLR-Bench-German is the German-language pendant of the original SLR-Benchdataset. It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into German. This enables systematic evaluation and training of Large Language Models (LLMs) in logical… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-German.tabular10K<n<100K2 likes420 downloads4mo agoHugging Face24kangaroo-dataset-german /kangaroo_dataset German Kangaroo Benchmark The complete German Mathematical Kangaroo archive from 1998 to 2025 as a multiple-choice benchmark: 3,886 items from 140 exams in five grade groups (3--4, 5--6, 7--8, 9--10, 11--13), worth 3, 4, or 5 points each. 1,746 items are multimodal, with a question diagram, image-based answer options, or both. The accompanying paper describes the extraction, the evaluation protocol, and the results. Files kangaroo.parquet: the benchmark, 3,886… See the full description on the dataset page: https://huggingface.co/datasets/kangaroo-dataset-german/kangaroo_dataset.textquestion-answering1K<n<10K0 likes418 downloads8d agoHugging Face25datastuff /german_maudio100K<n<1M0 likes415 downloads1y agoHugging Face26mteb /germanquad-retrieval GermanQuAD-Retrieval An MTEB dataset Massive Text Embedding Benchmark Context Retrieval for German Question Answering Task category t2t Domains Written, Non-fiction, Web Reference https://huggingface.co/datasets/deepset/germanquad How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["GermanQuAD-Retrieval"]) evaluator = mteb.MTEB(task) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/germanquad-retrieval.texttext-retrieval1K<n<10K0 likes381 downloads1y agoHugging Face27Muckylixx /german-sohee-synthetic-tts-24k Flevi Restiti Vici Bis uns der Weg nur noch nach vorn blieb This repository contains a synthetic German audiobook and text-to-speech training dataset based on the original novel: Flevi Restiti ViciBis uns der Weg nur noch nach vorn blieb The novel was written in German by Maurice Hartmann, who is the author and copyright holder. Work in progress This dataset is a work in progress. Future revisions may include corrected transcripts, regenerated audio… See the full description on the dataset page: https://huggingface.co/datasets/Muckylixx/german-sohee-synthetic-tts-24k.audiotext-to-speech1K<n<10K0 likes369 downloads3mo agoHugging Face28sprinklr-huggingface /CXM_Arena_German Dataset Card for CXM Arena German Benchmark Suite Dataset Description This dataset, "CXM Arena German Benchmark Suite," is a comprehensive collection designed to evaluate various AI capabilities within the Customer Experience Management (CXM) domain, specifically for the German language. It is closely modeled after the original CXM_Arena benchmark, but all data is in German. The suite consolidates five distinct tasks into a unified benchmark, enabling robust testing of… See the full description on the dataset page: https://huggingface.co/datasets/sprinklr-huggingface/CXM_Arena_German.document1K<n<10K0 likes356 downloads1y agoHugging Face29keremberke /german-traffic-sign-detection Dataset Labels ['animals', 'construction', 'cycles crossing', 'danger', 'no entry', 'pedestrian crossing', 'school crossing', 'snow', 'stop', 'bend', 'bend left', 'bend right', 'give way', 'go left', 'go left or straight', 'go right', 'go right or straight', 'go straight', 'keep left', 'keep right', 'no overtaking', 'no overtaking -trucks-', 'no traffic both ways', 'no trucks', 'priority at next intersection', 'priority road', 'restriction ends', 'restriction ends -overtaking… See the full description on the dataset page: https://huggingface.co/datasets/keremberke/german-traffic-sign-detection.imageobject-detectionn<1K9 likes343 downloads4y agoHugging Face30rusheeliyer /german-courts Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/rusheeliyer/german-courts.text1K<n<10K1 likes339 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.