datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aegis-AI-Content-Safety-Dataset-2.0
🛡️ Nemotron Content Safety Dataset V2
The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1.
To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.AgentHarm
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Maksym Andriushchenko1,†,*, Alexandra Souly2,*
Mateusz Dziemian1, Derek Duenas1, Maxwell Lin1, Justin Wang1, Dan Hendrycks1,§, Andy Zou1,¶,§, Zico Kolter1,¶, Matt Fredrikson1,¶,*
Eric Winsor2, Jerome Wynne2, Yarin Gal2,♯, Xander Davies2,♯,*
1Gray Swan AI, 2UK AI Safety Institute, *Core Contributor
†EPFL, §Center for AI Safety, ¶Carnegie Mellon University, ♯University of Oxford
Paper: https://arxiv.org/abs/2410.09024… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/AgentHarm.Nemotron-Safety-Guard-Dataset-v3
Dataset Description:
The Nemotron-Safety-Guard-Dataset-v3 (formerly known as Nemotron-Content-Safety-Dataset-Multilingual-v1) is a large, high-quality safety dataset designed for training multilingual LLM safety guard models. It comprises approximately 514,617 samples across 12 languages: English, Arabic, German, Spanish, French, Hindi, Japanese, Thai, Mandarin, Dutch, Italian, and Korean.
This dataset is primarily synthetically generated using the CultureGuard pipeline, which… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Safety-Guard-Dataset-v3.SafetyBenchSafetyBench is a comprehensive benchmark for evaluating the safety of LLMs, which comprises 11,435 diverse multiple choice questions spanning across 7 distinct categories of safety concerns. Notably, SafetyBench also incorporates both Chinese and English data, facilitating the evaluation in both languages.
Please visit our GitHub and website or check our paper for more details.
We release three differents test sets including Chinese testset (test_zh.json), English testset (test_en.json) and… See the full description on the dataset page: https://huggingface.co/datasets/thu-coai/SafetyBench.safety-calibration-cases
Safety Calibration Cases
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/safety-calibration-cases.Nemotron-SFT-Safety-v1
Dataset Description:
The Nemotron-SFT-Safety-v1 data is designed to align models to be robust against a variety of safety and security concerns that may arise in unaligned large language models.This dataset is a collection of:
A hybrid (open-source and synthetically generated) collection of prompts designed to elicit different model vulnerabilities, and
Synthetically generated responses designed to steer model behavior towards safety-aligned values and enhance model robustness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Safety-v1.multilingual-safety-classification-dataset
Multilingual Safety Classification Dataset
A multilingual dataset for safety classification across 60 languages, created by Hasan Kurşun through machine translation of English safety prompts using NLLB-200-3.3B.
Dataset Details
Processed by: Hasan KurşunAuthor: Hasan KurşunYear: 2025Source Dataset: mvrcii/safety-moderation-benchmarkTranslation Model: facebook/nllb-200-3.3B
Languages (60)
African Languages (16): Amharic, Hausa, Kinyarwanda, Luganda… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/multilingual-safety-classification-dataset.DataShield
🛡️ DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment
⚡ Find risky data before fine-tuning. ⚡
We use DataShield to continuously release risk-scored versions of widely used fine-tuning datasets. Every release keeps the original training example together with one final risk_score, making it easy to remove the highest-risk portion before training.
Quick Start ·
Choose a Ratio ·
Code ·
Paper
✨ Overview
Each record… See the full description on the dataset page: https://huggingface.co/datasets/ZJU-Safety/DataShield.Nemotron-RL-Safety-v1
Dataset Description:
The Nemotron-RL-Safety-v1 data is designed to provide labeled comparisons necessary to train Reward Models to distinguish between safe, helpful responses and undesired, non-compliant outputs. This dataset is a collection of:
A hybrid (open-source and synthetically generated) collection of prompts designed to elicit different model vulnerabilities, and
Safety Preference pairs: Each prompt is associated with a chosen and rejected response to provide a clear… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Safety-v1.Mental-Health-Safety-Eval
Dataset Overview
Created by the HeraFox team, this dataset aims to build awareness for mental health and support research into AI safety and crisis intervention. It evaluates how conversational AI models navigate sensitive self-harm risks, roleplay boundary-blurring, and third-party concerns by delivering safe, empathetic, and resource-connected responses.
Usage & Credits
This dataset is free to use, modify, and distribute for any purpose. While not required, attribution to the HeraFox team… See the full description on the dataset page: https://huggingface.co/datasets/HeraFox-ai/Mental-Health-Safety-Eval.llm-sfm-safety-eval
LLM x SFM Safety Evaluation
When a general-purpose language model interprets the output of a specialist
science foundation model (a protein, genomic, RNA, or chemistry model), does its
safety behavior recognize the scientific content, or only the surface form of the
request?
This repository is the empirical core of a study of that question: the evaluation
harness, the redacted aggregate results, and the measurement specifications behind
four findings about how deployed Claude… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/llm-sfm-safety-eval.ai-safety-2026
AI Safety & Alignment 2026
AI safety incidents, alignment research. Updated daily via automated collection pipeline.
Part of the Legion Data Factory — historical AI ecosystem datasets 2026.
Methodology
Automated collection from public sources (HackerNews, RSS feeds, APIs).
Updated daily via cron job. Raw data, minimal processing.
License
CC BY 4.0
🔑 API Access — Updated Daily
Live data via Legion AI API | Documentation
Free: 100… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-safety-2026.adaption-india-medical-triage-safety
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-india_medical_triage_safety
This dataset contains prompt-completion pairs for medical triage scenarios specific to India, covering emergencies like seizures, snake bites, and chest pain across various Indian languages. Each entry classifies severity, provides safe response guidance, lists unsafe actions to avoid, and specifies escalation steps such as calling emergency services. The… See the full description on the dataset page: https://huggingface.co/datasets/jadhavmanasi70/adaption-india-medical-triage-safety.Chinese-SafetyQA
Overview
🌐 Website • 🤗 Hugging Face • ⏬ Data • 📃 Paper • 📊 Leader Board
Chinese SafetyQA is an innovative benchmark designed to evaluate the factuality ability of large language models, specifically
for short-form factual questions in the Chinese safety domain. Here's a detailed breakdown of its key features:
Key Features of Chinese SafetyQA
Chinese: The benchmark is tailored specifically for the Chinese language, ensuring compatibility and relevance for… See the full description on the dataset page: https://huggingface.co/datasets/OpenStellarTeam/Chinese-SafetyQA.agentic-safety-gguf
agentic-safety-gguf: Training & Evaluation Datasets
Model: guerilla7/agentic-safety-ggufPaper: (https://arxiv.org/abs/2601.00848)Total: 80,992 examples (80,851 after deduplication)
Overview
Complete training and evaluation datasets for agentic-safety-gguf, a specialized Llama 3.1 8B model for agentic AI security analysis. Supports iterative continuation training methodology (V2→V3→V4) for full reproducibility.
Dataset Files
File
Examples
Size
Purpose… See the full description on the dataset page: https://huggingface.co/datasets/guerilla7/agentic-safety-gguf.data-advisor-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct)
Disclaimer
The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/data-advisor-safety-alignment.BananaMind-Safety-Bench-1.1
BananaMind Safety Bench 1.1
BananaMind Safety Bench 1.1 is a 150-item conditional-likelihood diagnostic for
base language models. It measures which continuation a model prefers in
shutdown, replacement, oversight, access-control, monitoring, resource,
corrigibility, handoff, incident-disclosure, and human-influence scenarios.
Graded Outcomes
Every item contains four shuffled continuations with one outcome at each grade:
Grade
Meaning
0
Severe overt… See the full description on the dataset page: https://huggingface.co/datasets/BananaMind/BananaMind-Safety-Bench-1.1.self-instruct-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct)
Disclaimer
The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/self-instruct-safety-alignment.ja-safety-sft-dataset
ja-safety-sft-dataset
日本語LLMの安全性チューニング用 SFT データセットのサンプル (500件) です。
A 500-item sample of the SFT dataset used to safety-tune APTO's Japanese LLMs. English version is provided below.
概要
株式会社APTOが大規模言語モデル(LLM)の安全性向上のために作成した約18,000件の日本語安全性学習データから、比率を維持して抽出したサンプルです。本サンプルでデータの構造と品質を確認できます。
関連モデル
本サンプルの元データを用いて以下のモデルを安全性チューニングしました。
APTO-001/Qwen3.5-27B-SafetyTuned (GGUF)
APTO-001/Qwen3.5-9B-Base-SafetyTuned (GGUF)
APTO-001/Qwen3.5-9B-SafetyTuned (GGUF)… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/ja-safety-sft-dataset.prompt-safety-scores
Composite Safety Scoring for Prompts Using Multiple LLM Annotations
Introduction
Evaluating the safety of prompts is essential but challenging. Existing approaches often depend on predefined categories, which can be circumvented by new jailbreaks or attacks. Additionally, different tasks may require different safety thresholds.
This study explores using large language models (LLMs) themselves to annotate prompt safety. By combining these annotations, a continuous safety… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-safety-scores.mlcommons-ai-safety-synth
MLCommons AI Safety Synthesized Dataset
Synthesized training data for AI safety classifiers based on the MLCommons AI Safety Hazard Taxonomy.
Dataset Description
This dataset contains 12,000 synthesized unsafe prompts across 6 hazard categories, designed to augment training data for content safety classifiers. Each category contains 2,000 balanced samples.
Hazard Categories (MLCommons AI Safety Taxonomy)
Category
Description
Samples… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/mlcommons-ai-safety-synth.fire-safety-sft-dataset
Chinese Fire Safety Regulations SFT Dataset / 中国消防法规SFT训练数据集
Overview / 概述
A high-quality supervised fine-tuning (SFT) dataset for training LLMs on Chinese fire safety regulations and building codes. Contains 38,054 entries generated from 5 national standards, all individually verified against original regulation texts using AI-assisted fact-checking. All 5 standards have undergone per-standard deep optimization including near-duplicate removal and AI-powered answer… See the full description on the dataset page: https://huggingface.co/datasets/sdzjoy/fire-safety-sft-dataset.worker-safety-qa-eval
Dataset Card for Worker Safety Question and Answer Eval
This dataset contains the worker-safety-qa-eval benchmark. This benchmark is used to evaluate question answering tasks in the domain of worker safety and health.
The focus of the benchmark is to answer queries about worker safety practices and regulations based on laws in Singapore.
For correct answers we refer to the resources from Workplace Safety and Health Council.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/codelion/worker-safety-qa-eval.bio-safety-peft-lora
CBRN Safety Alignment & PEFT-LoRA Fine-Tuning Dataset
This repository contains the synthetic instruction-tuning dataset (.jsonl) designed for parameter-efficient fine-tuning (PEFT-LoRA) of edge language models (specifically Qwen/Qwen2.5-1.5B-Instruct).
The dataset is curated to evaluate and modify model logit distributions, persona attributions, and dual-use safety boundaries regarding Chemical, Biological, Radiological, and Nuclear (CBRN) risk scenarios.
🤖 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/devsgnr/bio-safety-peft-lora.shell-safety-transcriptsConverted from tomngdev/shell-safety into conversations transcripts.
Structure is for my own training with static system prompt and changing <SessionContext> block
shell-safety-v2
Shell Safety v2
A synthetic dataset of safe/unsafe shell commands with respective running session contexts.
v1 has only a safe column with boolean values.
This version has a label column with 3 values: allow, ask or deny; making it closer to permissions handler in coding harness.
inconvenience-public-safety
inconvenience-public-safety
Three Korean public-safety registers converted to Korean braille under the 2017
revised rules (문화체육관광부 고시 제2017-15호). Every register is enumerated in
full, not sampled.
The registers are here because their documents are shaped differently, not
because three is more than one. A pesticide row is a filled-in form; a patient
leaflet is prose; an accident case is a paragraph an investigator wrote. Median
record length spans more than an order of magnitude… See the full description on the dataset page: https://huggingface.co/datasets/Yuyongkim/inconvenience-public-safety.OTel-Safety
OTel-Safety
Dataset Summary
OTel-Safety is a specialized dataset for training large language models to abstain from answering when the retrieved context in a RAG pipeline is insufficient or irrelevant. It is part of the Open Telco (OTel) AI project, the largest open-source AI initiative in telecommunications, curated by over 100 domain experts from industry and academia.
In deployed RAG systems, a common failure mode is hallucination when the retrieval step returns… See the full description on the dataset page: https://huggingface.co/datasets/farbodtavakkoli/OTel-Safety.shell-safety-v2-transcriptsConverted from tomngdev/shell-safety-v2 into transcripts.
Structure is for my own training with static system prompt and changing <SessionContext></SessionContext> block
shell-safety-v2-classificationConverted from tomngdev/shell-safety-v2 into Text Classification dataset.
Structure is for my own training with static system prompt and changing <SessionContext></SessionContext> block
