datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Data-Analytics-Digital-Marketing-Project-Management-QA_DBproduct-management-sft-100k
Product Management SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level product management across PRD writing, feature prioritization, OKR setting, roadmap planning, user research, competitive analysis, and stakeholder communication.
Motivation
AI assistants for product management commonly fail by:
Generic frameworks without application: Explaining RICE scoring without actually scoring the user's features; describing OKRs without writing them… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/product-management-sft-100k.task719_mmmlu_answer_generation_management
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task719_mmmlu_answer_generation_management
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task719_mmmlu_answer_generation_management.DoD-Instruction-6055-17-Emergency-Management-Program
🚨 DoD Emergency Management Program
Maintainer: Terry Eppler
Owner: US Federal Government
Source: DoD Instruction 6055.17
Source Version: Change 4, effective December 1, 2025
Ownership of Source: United States Department of Defense
📋 Overview
Dataset Summary
The DoD Emergency Management Program Question-Answer Dataset is a structured, document-grounded natural-language dataset derived from DoD Instruction 6055.17, “DoD Emergency… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-6055-17-Emergency-Management-Program.tool-reasoning-sft-TOOLS-context-management-handling
Tool Reasoning SFT — Context Management
A mixed-domain tool-use SFT dataset for training context-aware reasoning with structured tool interactions.
Format
Each row contains a JSON-serialized message list following a multi-role conversation format with tool definitions and calls.
Usage
from datasets import load_dataset
ds = load_dataset("AmanPriyanshu/tool-reasoning-sft-TOOLS-context-management-handling", split="train")
License
Apache 2.0
Incident-Management-Handbook
FEMA Incident Management Handbook Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on the Federal Emergency Management Agency Incident Management Handbook.
The FEMA Incident Management Handbook is an operational reference for FEMA personnel assigned to incident-level response and recovery missions. It describes FEMA incident-management… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/Incident-Management-Handbook.DoD-Instruction-8170-01-Online-Information-Management-And-Electronic-Messaging
📚 DoD Instruction 8170.01 Online Information Management and Electronic Messaging
Maintainer: Terry Eppler
Ownership: U.S. Department of Defense
📋 Overview
Dataset Summary
The DoD Instruction 8170.01 Online Information Management and Electronic Messaging
Dataset is a structured natural-language question-answering dataset derived from
DoD Instruction 8170.01, Online Information Management and Electronic Messaging.
DoD Instruction 8170.01… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8170-01-Online-Information-Management-And-Electronic-Messaging.Program-Management-Improvement-Accountability-Act-of-2016
Dataset Description
Maintainer: Terry Eppler
Owner: US Federal Government
The **Program Management Improvement Accountability Act of 2016 ** is an English-language instructional dataset containing 150 substantive question-and-answer records derived from the Program Management Improvement Accountability Act of 2016.
The Act, commonly abbreviated as PMIAA, was enacted as Public Law 114-264 on December 14, 2016. It amended title 31 of the United States Code to strengthen… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/Program-Management-Improvement-Accountability-Act-of-2016.DOD-Instruction-1400-25-Civilian-Personnel-Management-System
DoD Civilian Performance Management Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains 250 document-grounded question-and-answer records based on DoD Instruction 1400.25, Volume 430, “DoD Civilian Personnel Management System: Performance Management,” dated January 17, 2025, and incorporating Change 1 effective January 27, 2026.
The source establishes the Department of Defense framework for… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Instruction-1400-25-Civilian-Personnel-Management-System.DOD-Directive-type-Memorandum-22-001-Records-Management-Standards
DoD Records Management Standards for IT Systems and Services
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on Department of Defense Directive-type Memorandum 22-001, “DoD Standards for Records Management Capabilities in Programs Including Information Technology,” dated March 3, 2022, and incorporating Change 2 effective February 22, 2024.
The source establishes… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Directive-type-Memorandum-22-001-Records-Management-Standards.DOD-Directive-8000-01-Management-Of-Defense-Information
DoD Directive 8000.01 Management of the DoD Information Enterprise Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on Department of Defense Directive 8000.01, “Management of the Department of Defense Information Enterprise,” dated March 17, 2016, and incorporating Change 1 effective July 27, 2017.
The directive establishes Department-wide… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Directive-8000-01-Management-Of-Defense-Information.smolified-medimind-personalized-medication-adherence-assistant-for-chronic-disease-management
🤏 smolified-medimind-personalized-medication-adherence-assistant-for-chronic-disease-management
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model bubun123/smolified-medimind-personalized-medication-adherence-assistant-for-chronic-disease-management.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 13675e9f)
Records: 926
Type: Synthetic Instruction Tuning Data
⚖️… See the full description on the dataset page: https://huggingface.co/datasets/bubun123/smolified-medimind-personalized-medication-adherence-assistant-for-chronic-disease-management.pcos-management-patient-qa-from-eshre-guideline
Dataset Card for pcos-management-patient-qa-from-eshre-guideline
Dataset Details
Dataset Description
pcos-management-patient-qa-from-eshre-guideline is a clinically grounded conversational dataset designed to support training and evaluation of chat-based AI models for patient education in Polycystic Ovary Syndrome (PCOS).
The dataset contains structured user–assistant conversations derived from evidence-based recommendations in the International Evidence-based… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/pcos-management-patient-qa-from-eshre-guideline.itsm-change-management-benchmark
ITSM Change Management Benchmark
The first public dataset for evaluating AI agents on IT Service Management (ITSM) tasks, specifically ITIL Change Management RFC generation.
Dataset Description
This dataset contains structured ITSM data across three realistic enterprise scenarios, designed to benchmark AI agents that generate or evaluate Request for Change (RFC) documents against ITIL v4 standards.
Scenarios
Scenario
Category
Incidents
CMDB Items
Risk… See the full description on the dataset page: https://huggingface.co/datasets/VuduVations/itsm-change-management-benchmark.Change_Management_1
Change Management 1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required licenses… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Change_Management_1.Performance_Management_Difficult_Conversations_Practical
Performance Management Difficult Conversations — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Performance_Management_Difficult_Conversations_Practical.Performance_Management_Difficult_Conversations_Theory
Performance Management Difficult Conversations — Theory
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Performance_Management_Difficult_Conversations_Theory.Prioritization_Time_Attention_Management_Practical
Prioritization Time Attention Management — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Prioritization_Time_Attention_Management_Practical.Pomodoro_Technique_for_Time_Management_Corpus
Pomodoro Technique for Time Management
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
relevance_score: Relevance to the subject (0-1)
quality_score: Content quality score (0-1)
topics: JSON array of detected topics
character_count: Length of… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Pomodoro_Technique_for_Time_Management_Corpus.Time_management_Corpus
Time Management
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
relevance_score: Relevance to the subject (0-1)
quality_score: Content quality score (0-1)
topics: JSON array of detected topics
character_count: Length of the text
subject_name:… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Time_management_Corpus.Prioritization_Time_Attention_Management_Theory
Prioritization Time Attention Management — Theory
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Prioritization_Time_Attention_Management_Theory.Stress_Management_Corpus
stress-management-content
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
relevance_score: Relevance to the subject (0-1)
quality_score: Content quality score (0-1)
topics: JSON array of detected topics
character_count: Length of the text… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Stress_Management_Corpus.Stakeholder_Management_Engagement_Theory
Stakeholder Management Engagement — Theory
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Stakeholder_Management_Engagement_Theory.Change_Management_2
Change Management 2
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required licenses… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Change_Management_2.Stakeholder_Management_Engagement_Practical
Stakeholder Management Engagement — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Stakeholder_Management_Engagement_Practical.
