datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
german-courts
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/rusheeliyer/german-courts.German_Names_Central_And_Eastern_EuropeThis dataset contains German exonyms for various places in modern day Poland, Czech Republic, Latvia, Lithuania and Estonia.
Exonym : - A placename that is used by people who are not locals. For example, Prague is the Eng. exonym of Czech capital Praha, or Cologne is an exonym for German city Köln.
Due to extensive historical German rule and presence over large chunks of modern day Poland and Czech republic, these two countries populate the dataset the most.
datos-propiedadeshatecheck-german
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-german.german-credit-risk_credit-scoring_mlp
🏦 German Credit Risk - Dataset para MLP
Este dataset es parte del curso de Deep Learning impartido en el canal de YouTube de inGeniia. Se utiliza para demostrar la implementación de un Perceptrón Multicapa (MLP) para tareas de clasificación binaria (riesgo crediticio).
Descripción del Proyecto
El objetivo de este dataset es predecir si un cliente representa un buen o mal riesgo crediticio basándose en una serie de atributos financieros y personales.
Problema:… See the full description on the dataset page: https://huggingface.co/datasets/inGeniia/german-credit-risk_credit-scoring_mlp.german
German
The German dataset from the UCI ML repository.
Dataset on loan grants to customers.
Configurations and tasks
Configuration
Task
Description
encoding
Encoding dictionary showing original values of encoded features.
loan
Binary classification
Has the loan request been accepted?
Usage
from datasets import load_dataset
dataset = load_dataset("mstz/german", "loan")["train"]
Features
Feature
Type… See the full description on the dataset page: https://huggingface.co/datasets/mstz/german.german-polish-paired-placenames
Dataset Summary
This dataset contains the German and Polish names for almost 10k places in Poland. It has been generated using this code.
Many of these names are related to each other. Some German names are literal translation of the Polish names, some are phonetic modifications while some are unrelated.
Dataset Creation
Source Data
German wiki page
german-parliament-speeches
German Parliament Speeches
This dataset contains speeches from the German parliament, derived from the Open Discourse Project (Harvard Dataverse).
Source
Data source:
Open Discourse ProjectHarvard DataverseDOI: 10.7910/DVN/FIKIBO
Original citation:
@data{DVN/FIKIBO_2020,
author = {Richter, Florian and Koch, Philipp and Franke, Oliver and Kraus, Jakob and Kuruc, Fabrizio and Thiem, Anja and Högerl, Judith and Heine, Stella and Schöps, Konstantin},
publisher = {Harvard… See the full description on the dataset page: https://huggingface.co/datasets/emilpartow/german-parliament-speeches.nanochat-german-eval-data
nanochat German: Evaluation Data
This repository hosts the translated evaluation data used for assessing a German nanochat model.
Background information: The original nanochat implementation by Andrej Karpathy uses the "Mosaic Eval Gauntlet" (version v0.3.0) benchmark. More information about this benchmark can be found in Mosaic's blog post and this paper.
To evaluate our German nanochat model, we translated several datasets to German using Gemini 2.5 Pro. While this translation… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/nanochat-german-eval-data.textq-german
TextQ-German
TextQ investigates how people perceive the quality of machine-generated German text and how these subjective judgments can be modeled automatically. We identified task-specific quality dimensions, quantified them through user ratings, and developed models that predict perceived quality for new generated texts.
TextQ-German is a dataset suite for studying the Quality of Experience (QoE) of machine-generated German text. It covers two Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/nphamdinh/textq-german.germany_pv_dataGermanLanguageTwitterAntisemitism
A German Language Labeled Dataset of Tweets
Gunther Jikeli, Sameer Karali, Daniel Miehling and Katharina Soemer
{gjikeli, skarali, damieh, ksoemer}@iu.edu
Description
Our dataset contains 8,048 German language tweets related to Jewish life from a four-year timespan.
The dataset consists of 18 samples of tweets with the keyword “Juden” or “Israel.” The samples are representative samples of all live tweets (at the time of sampling) with these keywords respectively over… See the full description on the dataset page: https://huggingface.co/datasets/ISCA-IUB/GermanLanguageTwitterAntisemitism.Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP
Sylheti-Bangla-English-Russian-German Parallel Corpus for NLP
Welcome to the first open-source multilingual parallel corpus for the Sylheti (syl) language, engineered by a native Linguistics student. This dataset bridges Sylheti with four major global high-resource languages spanning three distinct language families (Indo-Aryan, Germanic, and Slavic) to support Computational Linguistics (CL), Natural Language Processing (NLP) research, and Large Language Model (LLM) fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/fahim-ling/Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP.human-robot-conversation-german
Human-Robot Dataset
The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the German language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language processing, and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-german.tibetan-to-german-translation-datasetThis dataset consists of three columns, the first of which is a sentence or phrase in Tibetan, the second is the phonetic transliteration of the Tibetan, and the third is the German translation of the Tibetan.
The dataset was scraped from Lotsawa House and is released under the same license as the texts from which it is sourced.
The dataset is part of the larger MLotsawa project, the code repo for which can be found here.
synthetic-spam-detection-dataset-german
Tanaos Spam Detection German Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate spam detection systems — models that detect, classify, or filter unsolicited commercial advertisement, fraudulent messages, or other unwanted content in text form — in German.
Our german spam detection model, tanaos-spam-detection-german, was trained on this dataset.
Dataset Summary
The… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-spam-detection-dataset-german.german-czech-paired-placenames
Dataset Summary
This dataset contains the German and corresponding Czech names for almost 5k places in Czech Republic. It has been generated using this code.
Many of these names are related to each other. Some German names are literal translation of the Czech names (or maybe the other way around), some are phonetic modifications while some are unrelated.
Dataset Creation
Source Data
English wiki page containing German exonyms for places in Czech Republic
german-school-system
The German School System 2026/27
A machine-readable snapshot of the German school system for the 2026/27 school
year: 16 federal states, 64 school types, 380 dated entries from 19 official
sources, 2 grading systems, 65 final qualifications. Every dated entry is
tagged with the role it concerns (225 student, 86 teacher, 69 exchange student).
Generated 2026-06-26. Published by Audecius.
The Kultusministerkonferenz publishes the holiday calendar as PDFs — no CSV, no
feed, no API… See the full description on the dataset page: https://huggingface.co/datasets/audecius/german-school-system.german-hate-speech-superset
German Hate Speech Superset
This dataset is a superset (N=50,545) of posts annotated as hateful or not. It results from the preprocessing and merge of all available German hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior, that… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/german-hate-speech-superset.German-Speech-Dataset
🎧 German Speech Dataset
The German Speech Dataset is a high-quality speech audio dataset designed to provide structured and scalable audio data for advanced AI and machine learning systems. It includes 142 hours of audio data across 768 files, delivered in MP3 and WAV formats, with a total size of 327 MB. This carefully curated audio dataset ensures diverse and representative voice data, with 53% male and 47% female speakers, and a balanced age distribution ranging from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/German-Speech-Dataset.German_sentimentgerman-speech-recognition-dataset
German Speech Dataset for recognition task
Dataset comprises 431 hours of telephone dialogues in German, collected from 590+ native speakers across various topics and domains, achieving an impressive 95% sentence accuracy rate. It is designed for research in automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural language processing (NLP). - Get the data… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/german-speech-recognition-dataset.supermarket-germany-salesTopic-specific-genre-classification_german_historical-newspapers
Dataset Card for Topic-specific Genre Classification of German Historical Newspapers
This dataset was developed to train and evaluate topic-specific genre classification of German-language historical newspaper clippings.
Curated by: [Sarah Oberbichler]
Language(s) (NLP): [German]
License: [afl-3.0]
Uses
Evaluation of machine learning models for topic-specific classification of ocr-processed historical texts with varying quality levels.
Fine-tuning models on… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Topic-specific-genre-classification_german_historical-newspapers.german-cities-open-data
InfraNode German Cities Open-Data Snapshot
Ein reproduzierbarer, offen lizenzierter Querschnitt von Infrastruktur- und
Umweltdaten für 84+ deutsche Städte, erzeugt aus der öffentlichen
InfraNode-API. Eine Zeile je Stadt.
Inhalt
Bereich
Felder
Quelle
Stammdaten
slug, name_de, state, ags, wikidata_qid, lat, lon, base_population, base_area_km2
Wikidata (CC0)
Wetter
weather_temperature_c, weather_humidity, weather_condition
DWD (GeoNutzV)
Luftqualität… See the full description on the dataset page: https://huggingface.co/datasets/Khaledc83/german-cities-open-data.german-english-email-ticket-classification
Customer Support Tickets (Short Version)
This dataset is a simplified version of the Customer Support Tickets dataset.
Dataset Details:
The dataset includes combinations of the following columns:
type
queue
priority
language
Modifications:
Shortened Version: This version only includes the first three rows for each combination of the above columns (i.e., 'type', 'queue', 'priority', 'language').
This reduction makes the dataset smaller and more manageable… See the full description on the dataset page: https://huggingface.co/datasets/ale-dp/german-english-email-ticket-classification.german-speech-recognition-dataset
German Telephone Dialogues Dataset - 431 Hours
Dataset comprises 431 hours of high-quality audio recordings from 590+ native German speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating German speech recognition systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of telephone dialogues in German for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/german-speech-recognition-dataset.synthetic-guardrail-dataset-german
Tanaos Guardrail German Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate guardrail systems — models that detect, classify, or filter unsafe, harmful or potentially dangerous content — in German. It can be used to train moderation models or integrate LLM safety filters for applications like chatbots, content generation, and user-facing AI systems.
Our german guardrail model… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-guardrail-dataset-german.GermanSentimentBank
GermanSentimentBank
tags: Sentiment Analysis, Language Modeling, Natural Language Processing
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'GermanSentimentBank' dataset is a collection of German text excerpts from various sources such as online reviews, social media posts, and forum discussions. The purpose of this dataset is to provide a diverse set of samples for training and evaluating sentiment analysis models tailored… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/GermanSentimentBank.lexical-decision-germanThis dataset contains German words/sentences for lexical decision tests, which we created with wuggy.
If you use this dataset, please cite the following preprint:
If you use this dataset, please cite the following publication:
@inproceedings{bunzeck-etal-2025-construction,
title = "Do Construction Distributions Shape Formal Language Learning In {G}erman {B}aby{LM}s?",
author = "Bunzeck, Bastian and
Duran, Daniel and
Zarrie{\ss}, Sina",
editor = "Boleda, Gemma and… See the full description on the dataset page: https://huggingface.co/datasets/bbunzeck/lexical-decision-german.
