CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rusheeliyer /german-courts Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/rusheeliyer/german-courts.text1K<n<10K1 likes339 downloads3y agoHugging Face02DebasishDhal99 /German_Names_Central_And_Eastern_EuropeThis dataset contains German exonyms for various places in modern day Poland, Czech Republic, Latvia, Lithuania and Estonia. Exonym : - A placename that is used by people who are not locals. For example, Prague is the Eng. exonym of Czech capital Praha, or Cologne is an exonym for German city Köln. Due to extensive historical German rule and presence over large chunks of modern day Poland and Czech republic, these two countries populate the dataset the most. texttranslation10K<n<100K1 likes336 downloads3y agoHugging Face03germancozzo /datos-propiedadestabularn<1K0 likes201 downloads2d agoHugging Face04Paul /hatecheck-german Dataset Card for Multilingual HateCheck Dataset Description Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish. For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate. This allows for targeted diagnostic insights into model performance. For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-german.tabulartext-classification1K<n<10K8 likes144 downloads4y agoHugging Face05inGeniia /german-credit-risk_credit-scoring_mlp 🏦 German Credit Risk - Dataset para MLP Este dataset es parte del curso de Deep Learning impartido en el canal de YouTube de inGeniia. Se utiliza para demostrar la implementación de un Perceptrón Multicapa (MLP) para tareas de clasificación binaria (riesgo crediticio). Descripción del Proyecto El objetivo de este dataset es predecir si un cliente representa un buen o mal riesgo crediticio basándose en una serie de atributos financieros y personales. Problema:… See the full description on the dataset page: https://huggingface.co/datasets/inGeniia/german-credit-risk_credit-scoring_mlp.tabulartabular-classification1K<n<10K3 likes84 downloads10mo agoHugging Face06mstz /german German The German dataset from the UCI ML repository. Dataset on loan grants to customers. Configurations and tasks Configuration Task Description encoding Encoding dictionary showing original values of encoded features. loan Binary classification Has the loan request been accepted? Usage from datasets import load_dataset dataset = load_dataset("mstz/german", "loan")["train"] Features Feature Type… See the full description on the dataset page: https://huggingface.co/datasets/mstz/german.tabulartabular-classification1K<n<10K0 likes77 downloads1y agoHugging Face07DebasishDhal99 /german-polish-paired-placenames Dataset Summary This dataset contains the German and Polish names for almost 10k places in Poland. It has been generated using this code. Many of these names are related to each other. Some German names are literal translation of the Polish names, some are phonetic modifications while some are unrelated. Dataset Creation Source Data German wiki page texttranslation1K<n<10K0 likes72 downloads3y agoHugging Face08emilpartow /german-parliament-speeches German Parliament Speeches This dataset contains speeches from the German parliament, derived from the Open Discourse Project (Harvard Dataverse). Source Data source: Open Discourse ProjectHarvard DataverseDOI: 10.7910/DVN/FIKIBO Original citation: @data{DVN/FIKIBO_2020, author = {Richter, Florian and Koch, Philipp and Franke, Oliver and Kraus, Jakob and Kuruc, Fabrizio and Thiem, Anja and Högerl, Judith and Heine, Stella and Schöps, Konstantin}, publisher = {Harvard… See the full description on the dataset page: https://huggingface.co/datasets/emilpartow/german-parliament-speeches.tabulartext-classification100K<n<1M4 likes72 downloads1y agoHugging Face09stefan-it /nanochat-german-eval-data nanochat German: Evaluation Data This repository hosts the translated evaluation data used for assessing a German nanochat model. Background information: The original nanochat implementation by Andrej Karpathy uses the "Mosaic Eval Gauntlet" (version v0.3.0) benchmark. More information about this benchmark can be found in Mosaic's blog post and this paper. To evaluate our German nanochat model, we translated several datasets to German using Gemini 2.5 Pro. While this translation… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/nanochat-german-eval-data.tabularn<1K0 likes53 downloads11mo agoHugging Face10nphamdinh /textq-german TextQ-German       TextQ investigates how people perceive the quality of machine-generated German text and how these subjective judgments can be modeled automatically. We identified task-specific quality dimensions, quantified them through user ratings, and developed models that predict perceived quality for new generated texts. TextQ-German is a dataset suite for studying the Quality of Experience (QoE) of machine-generated German text. It covers two Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/nphamdinh/textq-german.tabularn<1K1 likes52 downloads1mo agoHugging Face11Shark26 /germany_pv_datatabular10K<n<100K0 likes48 downloads8mo agoHugging Face12ISCA-IUB /GermanLanguageTwitterAntisemitism A German Language Labeled Dataset of Tweets Gunther Jikeli, Sameer Karali, Daniel Miehling and Katharina Soemer {gjikeli, skarali, damieh, ksoemer}@iu.edu Description Our dataset contains 8,048 German language tweets related to Jewish life from a four-year timespan. The dataset consists of 18 samples of tweets with the keyword “Juden” or “Israel.” The samples are representative samples of all live tweets (at the time of sampling) with these keywords respectively over… See the full description on the dataset page: https://huggingface.co/datasets/ISCA-IUB/GermanLanguageTwitterAntisemitism.tabular1K<n<10K0 likes46 downloads3y agoHugging Face13fahim-ling /Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLPgated Sylheti-Bangla-English-Russian-German Parallel Corpus for NLP Welcome to the first open-source multilingual parallel corpus for the Sylheti (syl) language, engineered by a native Linguistics student. This dataset bridges Sylheti with four major global high-resource languages spanning three distinct language families (Indo-Aryan, Germanic, and Slavic) to support Computational Linguistics (CL), Natural Language Processing (NLP) research, and Large Language Model (LLM) fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/fahim-ling/Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP.texttranslationn<1K1 likes46 downloads23d agoHugging Face14UniDataPro /human-robot-conversation-german Human-Robot Dataset The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the German language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language processing, and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-german.audioautomatic-speech-recognitionn<1K1 likes43 downloads1mo agoHugging Face15billingsmoore /tibetan-to-german-translation-datasetThis dataset consists of three columns, the first of which is a sentence or phrase in Tibetan, the second is the phonetic transliteration of the Tibetan, and the third is the German translation of the Tibetan. The dataset was scraped from Lotsawa House and is released under the same license as the texts from which it is sourced. The dataset is part of the larger MLotsawa project, the code repo for which can be found here. texttranslation10K<n<100K0 likes38 downloads2y agoHugging Face16tanaos /synthetic-spam-detection-dataset-german Tanaos Spam Detection German Training Dataset This dataset was created synthetically by Tanaos with the Artifex Python library. The dataset is designed to train and evaluate spam detection systems — models that detect, classify, or filter unsolicited commercial advertisement, fraudulent messages, or other unwanted content in text form — in German. Our german spam detection model, tanaos-spam-detection-german, was trained on this dataset. Dataset Summary The… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-spam-detection-dataset-german.texttext-classification10K<n<100K0 likes36 downloads7mo agoHugging Face17DebasishDhal99 /german-czech-paired-placenames Dataset Summary This dataset contains the German and corresponding Czech names for almost 5k places in Czech Republic. It has been generated using this code. Many of these names are related to each other. Some German names are literal translation of the Czech names (or maybe the other way around), some are phonetic modifications while some are unrelated. Dataset Creation Source Data English wiki page containing German exonyms for places in Czech Republic texttranslation1K<n<10K0 likes30 downloads3y agoHugging Face18audecius /german-school-system The German School System 2026/27 A machine-readable snapshot of the German school system for the 2026/27 school year: 16 federal states, 64 school types, 380 dated entries from 19 official sources, 2 grading systems, 65 final qualifications. Every dated entry is tagged with the role it concerns (225 student, 86 teacher, 69 exchange student). Generated 2026-06-26. Published by Audecius. The Kultusministerkonferenz publishes the holiday calendar as PDFs — no CSV, no feed, no API… See the full description on the dataset page: https://huggingface.co/datasets/audecius/german-school-system.tabularn<1K0 likes30 downloads1mo agoHugging Face19manueltonneau /german-hate-speech-supersetgated German Hate Speech Superset This dataset is a superset (N=50,545) of posts annotated as hateful or not. It results from the preprocessing and merge of all available German hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that: are documented are publicly available focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior, that… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/german-hate-speech-superset.tabulartext-classification10K<n<100K6 likes29 downloads2y agoHugging Face20Speech-data /German-Speech-Dataset 🎧 German Speech Dataset The German Speech Dataset is a high-quality speech audio dataset designed to provide structured and scalable audio data for advanced AI and machine learning systems. It includes 142 hours of audio data across 768 files, delivered in MP3 and WAV formats, with a total size of 327 MB. This carefully curated audio dataset ensures diverse and representative voice data, with 53% male and 47% female speakers, and a balanced age distribution ranging from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/German-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes25 downloads6mo agoHugging Face21sepidmnorozy /German_sentimenttext1K<n<10K2 likes23 downloads4y agoHugging Face22UniDataPro /german-speech-recognition-dataset German Speech Dataset for recognition task Dataset comprises 431 hours of telephone dialogues in German, collected from 590+ native speakers across various topics and domains, achieving an impressive 95% sentence accuracy rate. It is designed for research in automatic speech recognition (ASR) systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural language processing (NLP). - Get the data… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/german-speech-recognition-dataset.textautomatic-speech-recognitionn<1K1 likes22 downloads1mo agoHugging Face23steve-lynn /supermarket-germany-salestabular1K<n<10K0 likes22 downloads3mo agoHugging Face24oberbics /Topic-specific-genre-classification_german_historical-newspapers Dataset Card for Topic-specific Genre Classification of German Historical Newspapers This dataset was developed to train and evaluate topic-specific genre classification of German-language historical newspaper clippings. Curated by: [Sarah Oberbichler] Language(s) (NLP): [German] License: [afl-3.0] Uses Evaluation of machine learning models for topic-specific classification of ocr-processed historical texts with varying quality levels. Fine-tuning models on… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Topic-specific-genre-classification_german_historical-newspapers.imagetext-classificationn<1K0 likes21 downloads2y agoHugging Face25Khaledc83 /german-cities-open-data InfraNode German Cities Open-Data Snapshot Ein reproduzierbarer, offen lizenzierter Querschnitt von Infrastruktur- und Umweltdaten für 84+ deutsche Städte, erzeugt aus der öffentlichen InfraNode-API. Eine Zeile je Stadt. Inhalt Bereich Felder Quelle Stammdaten slug, name_de, state, ags, wikidata_qid, lat, lon, base_population, base_area_km2 Wikidata (CC0) Wetter weather_temperature_c, weather_humidity, weather_condition DWD (GeoNutzV) Luftqualität… See the full description on the dataset page: https://huggingface.co/datasets/Khaledc83/german-cities-open-data.tabularn<1K0 likes19 downloads3mo agoHugging Face26ale-dp /german-english-email-ticket-classification Customer Support Tickets (Short Version) This dataset is a simplified version of the Customer Support Tickets dataset. Dataset Details: The dataset includes combinations of the following columns: type queue priority language Modifications: Shortened Version: This version only includes the first three rows for each combination of the above columns (i.e., 'type', 'queue', 'priority', 'language'). This reduction makes the dataset smaller and more manageable… See the full description on the dataset page: https://huggingface.co/datasets/ale-dp/german-english-email-ticket-classification.texttext-classificationn<1K0 likes18 downloads1y agoHugging Face27ud-nlp /german-speech-recognition-dataset German Telephone Dialogues Dataset - 431 Hours Dataset comprises 431 hours of high-quality audio recordings from 590+ native German speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating German speech recognition systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of telephone dialogues in German for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/german-speech-recognition-dataset.textautomatic-speech-recognitionn<1K0 likes18 downloads8mo agoHugging Face28tanaos /synthetic-guardrail-dataset-german Tanaos Guardrail German Training Dataset This dataset was created synthetically by Tanaos with the Artifex Python library. The dataset is designed to train and evaluate guardrail systems — models that detect, classify, or filter unsafe, harmful or potentially dangerous content — in German. It can be used to train moderation models or integrate LLM safety filters for applications like chatbots, content generation, and user-facing AI systems. Our german guardrail model… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-guardrail-dataset-german.texttext-classification10K<n<100K0 likes17 downloads8mo agoHugging Face29infinite-dataset-hub /GermanSentimentBank GermanSentimentBank tags: Sentiment Analysis, Language Modeling, Natural Language Processing Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'GermanSentimentBank' dataset is a collection of German text excerpts from various sources such as online reviews, social media posts, and forum discussions. The purpose of this dataset is to provide a diverse set of samples for training and evaluating sentiment analysis models tailored… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/GermanSentimentBank.textn<1K0 likes14 downloads2y agoHugging Face30bbunzeck /lexical-decision-germanThis dataset contains German words/sentences for lexical decision tests, which we created with wuggy. If you use this dataset, please cite the following preprint: If you use this dataset, please cite the following publication: @inproceedings{bunzeck-etal-2025-construction, title = "Do Construction Distributions Shape Formal Language Learning In {G}erman {B}aby{LM}s?", author = "Bunzeck, Bastian and Duran, Daniel and Zarrie{\ss}, Sina", editor = "Boleda, Gemma and… See the full description on the dataset page: https://huggingface.co/datasets/bbunzeck/lexical-decision-german.text1K<n<10K0 likes12 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.