datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
permanitai-framework
⚠️ LIVING WORK DOCUMENT — DRAFT STATE ⚠️
This dataset is part of the AUGMANITAI Compendium, a living research work document, continuously updated. Each entry is a priority anchor for terminological provenance — not a final reference. Errors, omissions and improvements are expected and explicitly part of the evolving methodology.
LEBENDES ARBEITSDOKUMENT — ENTWURFSSTADIUM. Laufend aktualisiert. Prioritäts-Anker, nicht finale Referenz.
Author: Andreas Ehstand · ORCID: 0009-0006-3773-7796 ·… See the full description on the dataset page: https://huggingface.co/datasets/AndreasEhstand/permanitai-framework.ethical-framework-UNESCO-Ethics-of-AI
Ethical AI Training Dataset
Introduction
UNESCO's Ethics of Artificial Intelligence, adopted by 193 Member States in November 2021, represents the first global framework for ethical AI development and deployment.
While regional initiatives like The Montréal Declaration for a Responsible Development of Artificial Intelligence emphasize community-driven governance, UNESCO's approach establishes comprehensive international standards through coordinated multi-stakeholder… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/ethical-framework-UNESCO-Ethics-of-AI.ethical-framework
1. Dataset Title
Ethical AI Decision-Making Training Data (Montreal Declaration Edition)
2. Overview
This dataset contains carefully crafted scenarios (instructions) and detailed responses illustrating step-by-step ethical reasoning aligned with the principles outlined in the Montreal Declaration for Responsible AI. Each entry poses a complex ethical challenge and provides a reasoned solution while referencing the specific principle(s) being tested.
These entries can… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/ethical-framework.NIST-CyberSecurity-Framework
# NIST Cybersecurity Framework 2.0 Question Answering Dataset
Dataset Summary
The NIST Cybersecurity Framework 2.0 Question Answering Dataset is a synthetic
instruction-style question-answering dataset derived from the NIST Cybersecurity
Framework (CSF) 2.0.
The dataset is designed to support training, fine-tuning, retrieval evaluation, and
domain-specific question-answering use cases related to cybersecurity risk management,
cybersecurity governance, enterprise risk management… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-CyberSecurity-Framework.grounded-behavior-framework-v1_5
Grounded Behavior Framework N1 v1.5
Dataset sintético em português europeu para treino e avaliação de respostas
fundamentadas num contexto fornecido. Cada exemplo contém um contexto, uma
pergunta e uma resposta curta que aparece literalmente no contexto.
Como carregar
from datasets import load_dataset
dataset = load_dataset("empgces/grounded-behavior-framework-v1_5")
print(dataset)
print(dataset["train"][0])
Splits
Split
Exemplos
Utilização… See the full description on the dataset page: https://huggingface.co/datasets/empgces/grounded-behavior-framework-v1_5.NIST-AI-Risk-Management-Framework
# NIST AI Risk Management Framework Question Answering Dataset
Dataset Summary
The NIST AI Risk Management Framework Question Answering Dataset is a synthetic
instruction-style question-answering dataset derived from the NIST Artificial
Intelligence Risk Management Framework (AI RMF 1.0).
The dataset is designed to support training, fine-tuning, retrieval evaluation, and
domain-specific question-answering use cases related to AI risk management,
trustworthy AI, responsible AI… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-AI-Risk-Management-Framework.NIST-Privacy-Framework
NIST Privacy Framework Dataset
Dataset Summary
This dataset contains question-and-answer records derived from NIST Privacy Framework: A Tool for Improving Privacy Through Enterprise Risk Management, Version 1.0, published by the National Institute of Standards and Technology on January 16, 2020.
The source document provides a voluntary, risk-based framework for helping organizations improve privacy through enterprise risk management. It is designed to support privacy… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-Privacy-Framework.grc-security-frameworks
GRC Security Frameworks Dataset
A comprehensive dataset for training AI models on Governance, Risk, and Compliance (GRC) frameworks and cybersecurity standards.
Dataset Overview
This dataset contains 3,225 high-quality training examples covering major security and compliance frameworks. It's designed for fine-tuning large language models to become expert GRC assistants.
Covered Frameworks
CIS Controls v8.1.2 - 153 safeguards across 18 control families
Cloud… See the full description on the dataset page: https://huggingface.co/datasets/Zeezhu/grc-security-frameworks.dd-framework
📋 Due Diligence Framework
Core methodology, checklists, and templates for AI-powered due diligence analysis
This repository contains the foundational framework components for systematic due diligence analysis, including comprehensive checklists, structured question templates, and strategic analysis methodologies.
🎯 What's Included
📑 Due Diligence Checklists (2 files)
Comprehensive checklists covering all aspects of M&A due diligence:
original.md: 244 lines… See the full description on the dataset page: https://huggingface.co/datasets/jmzlx/dd-framework.qxf2-test-auto-framework-convs
Qxf2 Multi-Turn QA Dataset
Dataset Summary
The Qxf2 Multi-Turn QA Dataset is a collection of multi-turn question-answering (QA) conversations centered around Qxf2’s test automation framework. This dataset is designed to facilitate research and development in natural language understanding (NLU), conversational AI, and automation framework knowledge extraction.
Dataset Details
Domain: Software Testing, Test Automation
Purpose: Train and evaluate… See the full description on the dataset page: https://huggingface.co/datasets/shivaharip/qxf2-test-auto-framework-convs.sft-mobile-query-framework
SFT Mobile Query Framework Dataset
Dataset Description
This dataset contains 3607 training pairs for supervised fine-tuning (SFT) of language models to parse natural language queries about mobile phones into structured JSON execution plans.
Dataset Summary
Total Examples: 3607
Format: JSONL (question-answer pairs)
Task: Query Parsing & Structured Output Generation
Domain: Mobile Phone Specifications
Language: English
Purpose
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/sujitpandey/sft-mobile-query-framework.amber-framework-knowledge-pack
Amber Framework Knowledge Pack (demo)
The demo knowledge pack dataset behind
AgentC-Consulting/knowledge-packs:
teach a small local model the Amber web framework (Crystal),
and measure whether it learned anything with a before/after eval harness.
A knowledge pack compiles a body of expertise into curated sources, schema-validated
generated training JSONL, a contamination-guarded held-out eval set, and a JSON manifest.
This repo ships the exact training data and the two 50-item… See the full description on the dataset page: https://huggingface.co/datasets/crimson-knight/amber-framework-knowledge-pack.
