datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Romansh_German_Parallel_Data
Romansh–German Parallel Dataset (FineWeb-Based)
This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction.
Description
This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.Tab-MIA
Tab-MIA: A Benchmark for Membership Inference Attacks on Tabular Data
Tab-MIA is a benchmark dataset designed to evaluate the privacy risks of fine-tuning large language models (LLMs) on structured tabular data. It enables reproducible and systematic testing of Membership Inference Attacks (MIAs) across diverse datasets and six different serialization formats.
📋 Overview
Datasets:
WTQ (WikiTableQuestions)
WikiSQL
TabFact
Adult Census
California Housing… See the full description on the dataset page: https://huggingface.co/datasets/germane/Tab-MIA.german-wikipedia-clean-2german-public-sector
Dataset Card for public_sector_QA
Dataset Summary
public_sector_QA is a German-language question answering dataset focused on public sector and administrative-domain content. The file contains curated QA samples with source context and LLM-as-a-judge evaluation metadata.
This dataset was created from top-20-percent filtered outputs of three evaluated source files:
anlage-5-handbuch-offene-verwaltungsdaten_1_evaluated_top_20_percent.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Anirbanbhk/german-public-sector.german-public-sector-c_dbr
Dataset Card for public_sector_c_dbr_QA
Dataset Summary
public_sector_c_dbr_QA is a German-language QA dataset for public-sector and legal-administrative content. The dataset includes question, answer, source context, and LLM-as-a-judge quality metadata.
Dataset Files
Main file: public_sector_c_dbr_QA.jsonl
Base evaluated file: c_dbr_evaluated_top_20_percent.jsonl
Suggested split mapping: train only
Record Counts… See the full description on the dataset page: https://huggingface.co/datasets/Anirbanbhk/german-public-sector-c_dbr.Nekochu__Llama-3.1-8B-German-ORPO-details
Dataset Card for Evaluation run of Nekochu/Llama-3.1-8B-German-ORPO
Dataset automatically created during the evaluation run of model Nekochu/Llama-3.1-8B-German-ORPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nekochu__Llama-3.1-8B-German-ORPO-details.german-foerderprogramme-2026
Deutsche Weiterbildungsförderungen (berichtigte Fassung, Stand 19.09.2026)
Änderungsvermerk (19.09.2026): berichtigte Fassung
Diese Fassung ersetzt die Fassung vom 28.06.2026. Berichtigt wurden:
Bildungsgutschein: Seit dem 01.01.2025 stellt die Agentur für Arbeit den Bildungsgutschein auch für Leistungsberechtigte nach dem SGB II aus (§ 66a SGB II; § 16 Abs. 1 Satz 2 Nr. 4 SGB II ist weggefallen). Die Vorfassung nannte § 16 SGB II als Rechtsgrundlage und das… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/german-foerderprogramme-2026.german-court-decisions
Dataset Card for german-court-decisions
60k judicial decisions in Germany retrieved on January 1, 2024.
Dataset Description
Language(s) (NLP): German
License: MIT
Copyright notice: Automated retrieval of decisions from federal and state databases in Germany is permitted for non-commercial purposes only. As a result, the use of this dataset is permitted for non-commercial purposes only.
Uses
Prediction of verdicts based on statement of facts.
Direct Use… See the full description on the dataset page: https://huggingface.co/datasets/SH108/german-court-decisions.germanodia-german-parallel-corpus-research
Dataset Summary
This dataset is a high-quality, parallel corpus for Odia (Oriya) to German and German to Odia machine translation. It focuses on the news domain, specifically covering National, International, Sports, Trade, and Science & Technology topics.
The dataset contains 3,676 unique parallel sentence pairs, curated through a hybrid approach combining automated web scraping, manual human translation (Gold Standard), and human-corrected machine translation (Silver… See the full description on the dataset page: https://huggingface.co/datasets/abhinandansamal/odia-german-parallel-corpus-research.
