datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
un-digital-library
United Nations Digital Library (UNDL) Comprehensive Master Dataset
1. Executive Summary
Welcome to the United Nations Digital Library (UNDL) Comprehensive Master Dataset repository. This dataset represents a monumental effort to harvest, normalize, enrich, and democratize access to the vast archives of the United Nations. By leveraging advanced web harvesting techniques, robust state management, and modern big-data formats, this repository provides researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/un-digital-library.ouroboros-trace-help
Trace Help — does an execution trace help a model answer questions about a run?
In one minute. Twelve small programs in six languages (Python, JavaScript, C,
C++, Go, Elixir). Each was run once with a fixed command. Five questions per
program ask what actually happened on that one run: how many times a function
was called, what a particular call returned, what it was called with, whether a
function ran at all, which function raised. Sixty questions in total.
Every record carries… See the full description on the dataset page: https://huggingface.co/datasets/digitable-lol/ouroboros-trace-help.digitalisierungsmanager-curriculum-azav-2026
Digitalisierungsmanager für Prozessautomatisierung und Künstliche Intelligenz: Curriculum und AZAV-Zulassung
Änderungsvermerk (19.09.2026): berichtigte Fassung
Diese Fassung ersetzt die Fassung vom 25.05.2026. Berichtigt wurden:
Module und Unterrichtseinheiten: Modultitel und UE je Modul stehen jetzt im Wortlaut der AZAV-Zulassung (13 Module, zusammen 720 UE). Die Vorfassung enthielt Titel und eine UE-Verteilung, die es in der Zulassung nicht gibt, sowie eine… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/digitalisierungsmanager-curriculum-azav-2026.integreat-qa
Dataset
Our dataset consists of 906 diverse QA pairs in German and English.
The dataset is extractive, i.e., answers are given as sentence indices (breaking at the newline character \n).
Questions are automatically generated using an LLM.
The answers are manually annotated using voluntary crowdsourcing.
Repository: More Information Needed
Paper:
https://arxiv.org/abs/1806.03822
https://aclanthology.org/2024.konvens-main.25/
Our dataset is licensed under cc-by-4.0.
Properties… See the full description on the dataset page: https://huggingface.co/datasets/digitalfabrik/integreat-qa.
