datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Data-Analytics-Digital-Marketing-Project-Management-QA_DBproduct-management-sft-100k
Product Management SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level product management across PRD writing, feature prioritization, OKR setting, roadmap planning, user research, competitive analysis, and stakeholder communication.
Motivation
AI assistants for product management commonly fail by:
Generic frameworks without application: Explaining RICE scoring without actually scoring the user's features; describing OKRs without writing them… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/product-management-sft-100k.task719_mmmlu_answer_generation_management
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task719_mmmlu_answer_generation_management
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task719_mmmlu_answer_generation_management.DoD-Instruction-6055-17-Emergency-Management-Program
🚨 DoD Emergency Management Program
Maintainer: Terry Eppler
Owner: US Federal Government
Source: DoD Instruction 6055.17
Source Version: Change 4, effective December 1, 2025
Ownership of Source: United States Department of Defense
📋 Overview
Dataset Summary
The DoD Emergency Management Program Question-Answer Dataset is a structured, document-grounded natural-language dataset derived from DoD Instruction 6055.17, “DoD Emergency… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-6055-17-Emergency-Management-Program.tool-reasoning-sft-TOOLS-context-management-handling
Tool Reasoning SFT — Context Management
A mixed-domain tool-use SFT dataset for training context-aware reasoning with structured tool interactions.
Format
Each row contains a JSON-serialized message list following a multi-role conversation format with tool definitions and calls.
Usage
from datasets import load_dataset
ds = load_dataset("AmanPriyanshu/tool-reasoning-sft-TOOLS-context-management-handling", split="train")
License
Apache 2.0
Incident-Management-Handbook
FEMA Incident Management Handbook Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on the Federal Emergency Management Agency Incident Management Handbook.
The FEMA Incident Management Handbook is an operational reference for FEMA personnel assigned to incident-level response and recovery missions. It describes FEMA incident-management… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/Incident-Management-Handbook.DoD-Instruction-8170-01-Online-Information-Management-And-Electronic-Messaging
📚 DoD Instruction 8170.01 Online Information Management and Electronic Messaging
Maintainer: Terry Eppler
Ownership: U.S. Department of Defense
📋 Overview
Dataset Summary
The DoD Instruction 8170.01 Online Information Management and Electronic Messaging
Dataset is a structured natural-language question-answering dataset derived from
DoD Instruction 8170.01, Online Information Management and Electronic Messaging.
DoD Instruction 8170.01… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8170-01-Online-Information-Management-And-Electronic-Messaging.Program-Management-Improvement-Accountability-Act-of-2016
Dataset Description
Maintainer: Terry Eppler
Owner: US Federal Government
The **Program Management Improvement Accountability Act of 2016 ** is an English-language instructional dataset containing 150 substantive question-and-answer records derived from the Program Management Improvement Accountability Act of 2016.
The Act, commonly abbreviated as PMIAA, was enacted as Public Law 114-264 on December 14, 2016. It amended title 31 of the United States Code to strengthen… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/Program-Management-Improvement-Accountability-Act-of-2016.DOD-Instruction-1400-25-Civilian-Personnel-Management-System
DoD Civilian Performance Management Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains 250 document-grounded question-and-answer records based on DoD Instruction 1400.25, Volume 430, “DoD Civilian Personnel Management System: Performance Management,” dated January 17, 2025, and incorporating Change 1 effective January 27, 2026.
The source establishes the Department of Defense framework for… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Instruction-1400-25-Civilian-Personnel-Management-System.DOD-Directive-type-Memorandum-22-001-Records-Management-Standards
DoD Records Management Standards for IT Systems and Services
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on Department of Defense Directive-type Memorandum 22-001, “DoD Standards for Records Management Capabilities in Programs Including Information Technology,” dated March 3, 2022, and incorporating Change 2 effective February 22, 2024.
The source establishes… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Directive-type-Memorandum-22-001-Records-Management-Standards.DOD-Directive-8000-01-Management-Of-Defense-Information
DoD Directive 8000.01 Management of the DoD Information Enterprise Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on Department of Defense Directive 8000.01, “Management of the Department of Defense Information Enterprise,” dated March 17, 2016, and incorporating Change 1 effective July 27, 2017.
The directive establishes Department-wide… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Directive-8000-01-Management-Of-Defense-Information.smolified-medimind-personalized-medication-adherence-assistant-for-chronic-disease-management
🤏 smolified-medimind-personalized-medication-adherence-assistant-for-chronic-disease-management
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model bubun123/smolified-medimind-personalized-medication-adherence-assistant-for-chronic-disease-management.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 13675e9f)
Records: 926
Type: Synthetic Instruction Tuning Data
⚖️… See the full description on the dataset page: https://huggingface.co/datasets/bubun123/smolified-medimind-personalized-medication-adherence-assistant-for-chronic-disease-management.pcos-management-patient-qa-from-eshre-guideline
Dataset Card for pcos-management-patient-qa-from-eshre-guideline
Dataset Details
Dataset Description
pcos-management-patient-qa-from-eshre-guideline is a clinically grounded conversational dataset designed to support training and evaluation of chat-based AI models for patient education in Polycystic Ovary Syndrome (PCOS).
The dataset contains structured user–assistant conversations derived from evidence-based recommendations in the International Evidence-based… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/pcos-management-patient-qa-from-eshre-guideline.Change_Management_1
Change Management 1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required licenses… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Change_Management_1.Performance_Management_Difficult_Conversations_Practical
Performance Management Difficult Conversations — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Performance_Management_Difficult_Conversations_Practical.Performance_Management_Difficult_Conversations_Theory
Performance Management Difficult Conversations — Theory
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Performance_Management_Difficult_Conversations_Theory.Prioritization_Time_Attention_Management_Practical
Prioritization Time Attention Management — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Prioritization_Time_Attention_Management_Practical.Pomodoro_Technique_for_Time_Management_Corpus
Pomodoro Technique for Time Management
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
relevance_score: Relevance to the subject (0-1)
quality_score: Content quality score (0-1)
topics: JSON array of detected topics
character_count: Length of… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Pomodoro_Technique_for_Time_Management_Corpus.Time_management_Corpus
Time Management
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
relevance_score: Relevance to the subject (0-1)
quality_score: Content quality score (0-1)
topics: JSON array of detected topics
character_count: Length of the text
subject_name:… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Time_management_Corpus.Prioritization_Time_Attention_Management_Theory
Prioritization Time Attention Management — Theory
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Prioritization_Time_Attention_Management_Theory.Stress_Management_Corpus
stress-management-content
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
relevance_score: Relevance to the subject (0-1)
quality_score: Content quality score (0-1)
topics: JSON array of detected topics
character_count: Length of the text… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Stress_Management_Corpus.Stakeholder_Management_Engagement_Theory
Stakeholder Management Engagement — Theory
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Stakeholder_Management_Engagement_Theory.Change_Management_2
Change Management 2
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required licenses… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Change_Management_2.Stakeholder_Management_Engagement_Practical
Stakeholder Management Engagement — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Stakeholder_Management_Engagement_Practical.
