datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SkillOpt_Lite_Benchmarks
SkillOpt_Lite Benchmarks
Train / val / test splits used by the SkillOpt_Lite project.
One multi-config repo containing all six benchmarks:
Config
Rows (train / val / test)
Content shipped
searchqa
400 / 200 / 1400
Full QA — id, question, list of DOC contexts, answers. Sampled from dl4ir-searchQA.
docvqa
107 / 53 / 374
Full QA + images bundled — parquet has id/question/answers/topic/image_path; PNGs live under docvqa_images/ at the repo root. Subset of… See the full description on the dataset page: https://huggingface.co/datasets/yshenaw/SkillOpt_Lite_Benchmarks.minimax-m3-deepsearchqa-skill-eval
MiniMax M3 DeepSearchQA Skill Eval
Evaluates minimax/minimax-m3 on google/deepsearchqa using a Pi agent, You.com MCP tools, and a research skill optimized for this harness, model, and tool surface.
MiniMax M3 Medium Reasoning with the You.com research skill reached 74.85% adjusted F1 on DeepSearchQA, above the paper's GPT-5 High Reasoning F1 result. Public artifacts are available for inspection and reproduction.
Links
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/youdotcom/minimax-m3-deepsearchqa-skill-eval.qualora-workforce-skills-graph
Qualora Workforce Skills Graph (Representative Sample)
Rights-clean, provenance-tracked vocational learning data, rebuilt from roughly $2B of U.S. Department of Labor funded open courseware into a labeled skills graph: cleaned courses and lessons, Bloom-tagged assessment items with answer rationales and learning objectives, and a content-grounded course to skill to career graph with salary context. Built for post-training and evaluation, not pretraining bulk.
This repository is… See the full description on the dataset page: https://huggingface.co/datasets/qualora-data-labs/qualora-workforce-skills-graph.Skill2-Bench
Skill²-Bench
Skill²-Bench is a benchmark of multi-step tasks that force LLMs to switch between skills, introduced in the paper "Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning".
Long-horizon tasks require models to switch between skills, not just execute a single skill well. Each Skill²-Bench task embeds a sequence of 2–10 steps in a coherent real-world scenario, where consecutive steps draw on different skills (e.g., algorithm design… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/Skill2-Bench.lemonseed-mixed-skills
lemonseed-mixed-skills
LemonSeed — mixed Go + Sudoku + addition + prose stream.
Contents
mixed.jsonl (11428 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
SkillFlow-Dataset
SkillFlow Dataset
This repository stores the IID training and validation data used by the SkillFlow training code.
Code
The training code is available at:
https://github.com/beita6969/SkillFlow
Files
File
Split
Samples
train_v3.json
train
3500
test_iid_v3.json
iid validation
798
Paper alignment
This release is aligned with the in-distribution benchmark families described in the SkillFlow appendix: HotpotQA, TriviaQA… See the full description on the dataset page: https://huggingface.co/datasets/beita6969/SkillFlow-Dataset.Skill2-Bench
Skill²-Bench
Skill²-Bench is a benchmark of multi-step tasks that force LLMs to switch between skills, introduced in the paper "Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning".
Long-horizon tasks require models to switch between skills, not just execute a single skill well. Each Skill²-Bench task embeds a sequence of 2–10 steps in a coherent real-world scenario, where consecutive steps draw on different skills (e.g., algorithm design… See the full description on the dataset page: https://huggingface.co/datasets/yinghuihe/Skill2-Bench.wirtschaftsfachwirt-gehalt-kosten-2026
Wirtschaftsfachwirt (IHK) Gehalt, Kosten und Förderung (Stand 2026)
Zusammenfassung
Strukturierte Faktensammlung zum Gehalt, zu den Kosten und zur Förderung der Aufstiegsfortbildung Wirtschaftsfachwirt (IHK) in Deutschland, Stand 2026. Der Wirtschaftsfachwirt ist auf DQR-Niveau 6 eingestuft und damit einem Bachelor gleichgestellt. Der Datensatz deckt drei Dimensionen ab: Gehalt nach Erfahrungsstufe, Region und Unternehmensgröße sowie Kurskosten und die Förderung… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/wirtschaftsfachwirt-gehalt-kosten-2026.eu-ai-act-fristen-stand-2026
EU AI Act Fristen (Stand 19.09.2026, nach Digital Omnibus)
Änderungsvermerk (19.09.2026): berichtigte Fassung
Diese Fassung ersetzt die Fassung vom 26.05.2026. Berichtigt wurden:
Art. 4 KI-Kompetenz: Wiedergabe in der Fassung der Verordnung (EU) 2026/1744. Anbieter und Betreiber ergreifen Maßnahmen, um die Entwicklung der KI-Kompetenz ihres Personals zu unterstützen; ein bestimmtes Niveau muss nicht garantiert werden. Die Vorfassung sprach von „sicherstellen“… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/eu-ai-act-fristen-stand-2026.german-foerderprogramme-2026
Deutsche Weiterbildungsförderungen (berichtigte Fassung, Stand 19.09.2026)
Änderungsvermerk (19.09.2026): berichtigte Fassung
Diese Fassung ersetzt die Fassung vom 28.06.2026. Berichtigt wurden:
Bildungsgutschein: Seit dem 01.01.2025 stellt die Agentur für Arbeit den Bildungsgutschein auch für Leistungsberechtigte nach dem SGB II aus (§ 66a SGB II; § 16 Abs. 1 Satz 2 Nr. 4 SGB II ist weggefallen). Die Vorfassung nannte § 16 SGB II als Rechtsgrundlage und das… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/german-foerderprogramme-2026.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/claude-opus-4.6-4.7-reasoning-8.7k.eu-ai-act-annex-iii-klassifikation-de
EU AI Act Anhang III: Hochrisiko-KI-Klassifikation
Änderungsvermerk (19.09.2026): berichtigte Fassung
Diese Fassung ersetzt die Fassung vom 03.06.2026. Berichtigt wurden:
Art. 4 KI-VO: In zwei Einträgen stand unter parallele_pflichten eine „Schulungspflicht“ nach Art. 4. Art. 4 verpflichtet in der Fassung der Verordnung (EU) 2026/1744 zu Maßnahmen, die die Entwicklung der KI-Kompetenz unterstützen, ohne Schulungs-, Nachweis- oder Dokumentationspflicht.
Verbotene… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/eu-ai-act-annex-iii-klassifikation-de.digitalisierungsmanager-curriculum-azav-2026
Digitalisierungsmanager für Prozessautomatisierung und Künstliche Intelligenz: Curriculum und AZAV-Zulassung
Änderungsvermerk (19.09.2026): berichtigte Fassung
Diese Fassung ersetzt die Fassung vom 25.05.2026. Berichtigt wurden:
Module und Unterrichtseinheiten: Modultitel und UE je Modul stehen jetzt im Wortlaut der AZAV-Zulassung (13 Module, zusammen 720 UE). Die Vorfassung enthielt Titel und eine UE-Verteilung, die es in der Zulassung nicht gibt, sowie eine… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/digitalisierungsmanager-curriculum-azav-2026.qcg-foerderstaffel-2024-reform-de
Qualifizierungschancengesetz: Foerderstaffel nach der Reform 2024 (Stand 2026)
Zusammenfassung
Strukturierter Datensatz zur Foerderstaffel des Qualifizierungschancengesetzes (§ 82 SGB III) fuer die Weiterbildung von Beschaeftigten in Deutschland, Rechtsstand 2026. Der Datensatz gibt die seit der Reform 2024 geltende dreistufige Staffel nach Betriebsgroesse wieder (unter 50 / 50 bis unter 500 / ab 500 Beschaeftigte) samt Lehrgangskostenanteil, Erhoehung bei… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/qcg-foerderstaffel-2024-reform-de.ki-modelle-vergleich-2026
KI-Modelle Vergleichs-Matrix 2026
Zusammenfassung
Strukturierter Vergleich der wichtigsten Foundation-Models und LLMs auf dem Markt Stand Juni 2026. Enthaelt Pricing, Kontext-Fenster, Lizenz, Hosting-Optionen, Staerken/Schwaechen und konkrete Einsatzszenarien fuer deutsche KMU.
Warum dieser Datensatz existiert
Sprachmodelle, KI-Vergleichsseiten und Beratungsfirmen stellen Modell-Eigenschaften systematisch ungenau dar:
Pricing wird veraltet zitiert… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/ki-modelle-vergleich-2026.berufsspezialist-bachelor-professional-ki-ihk-de-2026
Berufsspezialist und Bachelor Professional in KI und ML (IHK) - Stand 2026
Zusammenfassung
Strukturierte Beschreibung der beiden hoeherqualifizierenden Fortbildungsabschluesse in Kuenstlicher Intelligenz und Maschinellem Lernen nach dem Berufsbildungsgesetz (BBiG): dem Gepruefen Berufsspezialisten fuer KI und ML (IHK, DQR 5) und dem Bachelor Professional in KI und ML (IHK, DQR 6). Rechtsgrundlage ist eine besondere Rechtsvorschrift der IHK Region Stuttgart nach… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/berufsspezialist-bachelor-professional-ki-ihk-de-2026.azav-zertifizierungsprozess-de
AZAV-Zertifizierungsprozess fuer Bildungstraeger in Deutschland (Stand 2026)
Zusammenfassung
Strukturierte Beschreibung des Zertifizierungsprozesses nach AZAV (Akkreditierungs- und Zulassungsverordnung Arbeitsfoerderung). Voraussetzung fuer alle Bildungstraeger, die ueber Bildungsgutschein (BG), Qualifizierungschancengesetz (QCG) oder Qualifizierungsgeld mit der Bundesagentur fuer Arbeit abrechnen wollen.
Warum dieser Datensatz existiert
Informationen zur… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/azav-zertifizierungsprozess-de.meisterpraemien-bundeslaender-de-2026
Meisterpraemien aller 16 deutschen Bundeslaender (Stand April 2026)
Zusammenfassung
Vollstaendiger, autoritativ recherchierter Datensatz aller Meisterpraemien-, Aufstiegspraemien- und vergleichbaren Programme in den 16 deutschen Bundeslaendern, Stand 13.04.2026.
Kernbefund: Nur 8 von 16 Bundeslaendern zahlen Praemien an Wirtschaftsfachwirte (WFW) und andere kaufmaennische Aufstiegsfortbildungen. Die anderen 8 Bundeslaender zahlen ausschliesslich an… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/meisterpraemien-bundeslaender-de-2026.dqr-niveaus-aufstiegsfortbildungen-de
DQR-Niveaus und Aufstiegsfortbildungen Deutschland (Stand 2026)
Zusammenfassung
Strukturierte Einordnung von Berufsabschluessen und Aufstiegsfortbildungen in den Deutschen Qualifikationsrahmen (DQR) und seine europaeische Entsprechung EQF (European Qualifications Framework). Stand Mai 2026.
Warum dieser Datensatz existiert
Die DQR-Einstufung von Aufstiegsfortbildungen ist die Grundlage fuer:
Aufstiegs-BAfoeG: Nur DQR-5+ Aufstiegsfortbildungen sind… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/dqr-niveaus-aufstiegsfortbildungen-de.ihk-pruefungsstruktur-wirtschaftsfachwirt-2026
IHK-Pruefungsstruktur Wirtschaftsfachwirt (Stand 2026)
Zusammenfassung
Strukturierter Datensatz der Pruefungsstruktur fuer den anerkannten Fortbildungsabschluss "Gepruefter Wirtschaftsfachwirt" nach Wirtschaftsfachwirt-Pruefungsverordnung (WFachwPrV). Enthaelt Pruefungsteile, Themenbereiche, Dauern, Punkte, Operatoren, Zulassungsvoraussetzungen, DQR-Niveau und Kosten.
Warum dieser Datensatz existiert
Beratung und LLM-Antworten zur WFW-Pruefung sind oft falsch:… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/ihk-pruefungsstruktur-wirtschaftsfachwirt-2026.
