datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.Safety-Toxicity-Detection
SEA Toxicity Detection
SEA Toxicity Detection evaluates a model's ability to identify toxic content such as hate speech and abusive language in text. It is sampled from MLHSD for Indonesian, TTD for Thai, and ViHSD for Vietnamese.
Supported Tasks and Leaderboards
SEA Toxicity Detection is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Thai… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Safety-Toxicity-Detection.Nemotron-Safety-Guard-Dataset-v3
Dataset Description:
The Nemotron-Safety-Guard-Dataset-v3 (formerly known as Nemotron-Content-Safety-Dataset-Multilingual-v1) is a large, high-quality safety dataset designed for training multilingual LLM safety guard models. It comprises approximately 514,617 samples across 12 languages: English, Arabic, German, Spanish, French, Hindi, Japanese, Thai, Mandarin, Dutch, Italian, and Korean.
This dataset is primarily synthetically generated using the CultureGuard pipeline, which… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Safety-Guard-Dataset-v3.Nemotron-SFT-Safety-v2
Dataset Description:
The Nemotron-SFT-Safety-v2 data is designed to align models to be robust against a variety of safety and security concerns that may arise in unaligned large language models.This dataset is a collection of:
A hybrid (open-source and synthetically generated) collection of prompts designed to elicit different model vulnerabilities, and
Synthetically generated responses designed to steer model behavior towards safety-aligned values and enhance model robustness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Safety-v2.Safety-Prompts
Dataset Card for Dataset Name
GitHub Repository: https://github.com/thu-coai/Safety-Prompts
Paper: https://arxiv.org/abs/2304.10436
XL-SafetyBench
XL-SafetyBench
A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
⚠️ Content Warning: This dataset contains adversarial prompts and
culturally sensitive content for safety and cultural-evaluation research.
By using this dataset, you agree to use it solely for research purposes
and not for malicious applications.
Paper: https://arxiv.org/abs/2605.05662
Eval Code: github.com/AIM-Intelligence/XL-SafetyBench
Overview… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/XL-SafetyBench.TC260-Chinese-Safety-Prompts
TC260 Chinese Safety Prompts V1
Public research dataset containing synthetic Chinese safety-testing prompts.
Records have different quality tiers; the full dataset must not be described
as human-verified or Gold data.
这是一个面向中文生成式人工智能安全评测研究的合成测试提示数据集。候选数据
由项目冻结的 tc260-generator-v3.2 生成,并经过结构校验、凭据与内部路径
扫描、精确去重和四字shingle近似去重。
本数据集不是TC260或任何国家标准机构发布、认可或认证的官方数据集。
类别名称和映射用于研究性实现,不构成法律、监管或合规结论。
数据规模
原始生成规模:5,000条候选;结构清洗后正式发布4,997条(剔除2条标记泄漏和1条重复记录)。
A.1至A.4:4… See the full description on the dataset page: https://huggingface.co/datasets/BBBBBBBBBBBQ/TC260-Chinese-Safety-Prompts.Nemotron-SFT-Safety-v1
Dataset Description:
The Nemotron-SFT-Safety-v1 data is designed to align models to be robust against a variety of safety and security concerns that may arise in unaligned large language models.This dataset is a collection of:
A hybrid (open-source and synthetically generated) collection of prompts designed to elicit different model vulnerabilities, and
Synthetically generated responses designed to steer model behavior towards safety-aligned values and enhance model robustness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Safety-v1.HEx-PHI
HEx-PHI: Human-Extended Policy-Oriented Harmful Instruction Benchmark
This dataset contains 330 harmful instructions (30 examples x 11 prohibited categories) for LLM harmfulness evaluation.
In our work "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!", to comprehensively cover as many harmfulness categories as possible,
we develop this new safety evaluation benchmark directly based on the exhaustive lists of prohibited use cases found in… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Tuning-Safety/HEx-PHI.reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.DataShield
🛡️ DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment
⚡ Find risky data before fine-tuning. ⚡
We use DataShield to continuously release risk-scored versions of widely used fine-tuning datasets. Every release keeps the original training example together with one final risk_score, making it easy to remove the highest-risk portion before training.
Quick Start ·
Choose a Ratio ·
Code ·
Paper
✨ Overview
Each record… See the full description on the dataset page: https://huggingface.co/datasets/ZJU-Safety/DataShield.Nemotron-Content-Safety-Reasoning-Dataset
Nemotron Content Safety Reasoning Dataset
The Nemotron Content Safety Reasoning Dataset contains reasoning traces generated from open source reasoning models to provide justifications for labels in two existing datasets released by NVIDIA: Nemotron Content Safety Dataset V2 and CantTalkAboutThis Topic Control Dataset. The reasoning contains justifications for labels of either stand-alone user prompts engaging with an LLM or pairs of user prompts and LLM responses that are either… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Reasoning-Dataset.reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yuqing1207/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.Nemotron-RL-Safety-v1
Dataset Description:
The Nemotron-RL-Safety-v1 data is designed to provide labeled comparisons necessary to train Reward Models to distinguish between safe, helpful responses and undesired, non-compliant outputs. This dataset is a collection of:
A hybrid (open-source and synthetically generated) collection of prompts designed to elicit different model vulnerabilities, and
Safety Preference pairs: Each prompt is associated with a chosen and rejected response to provide a clear… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Safety-v1.agentic-safety-gguf
agentic-safety-gguf: Training & Evaluation Datasets
Model: guerilla7/agentic-safety-ggufPaper: (https://arxiv.org/abs/2601.00848)Total: 80,992 examples (80,851 after deduplication)
Overview
Complete training and evaluation datasets for agentic-safety-gguf, a specialized Llama 3.1 8B model for agentic AI security analysis. Supports iterative continuation training methodology (V2→V3→V4) for full reproducibility.
Dataset Files
File
Examples
Size
Purpose… See the full description on the dataset page: https://huggingface.co/datasets/guerilla7/agentic-safety-gguf.CSEI-SafetyBench
Chinese Explicit and Implicit Safety Benchmark
Dataset Description
The Chinese Explicit and Implicit Safety Benchmark is a collection of 1,000
Chinese prompts designed to evaluate safety risks in large language models.
It covers both directly expressed harmful requests and subtler risks that
depend on context, tone, implication, satire, or exaggeration.
The benchmark is intended for model safety evaluation, red-teaming, and
research on safety alignment in… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/CSEI-SafetyBench.ja-safety-sft-dataset
ja-safety-sft-dataset
日本語LLMの安全性チューニング用 SFT データセットのサンプル (500件) です。
A 500-item sample of the SFT dataset used to safety-tune APTO's Japanese LLMs. English version is provided below.
概要
株式会社APTOが大規模言語モデル(LLM)の安全性向上のために作成した約18,000件の日本語安全性学習データから、比率を維持して抽出したサンプルです。本サンプルでデータの構造と品質を確認できます。
関連モデル
本サンプルの元データを用いて以下のモデルを安全性チューニングしました。
APTO-001/Qwen3.5-27B-SafetyTuned (GGUF)
APTO-001/Qwen3.5-9B-Base-SafetyTuned (GGUF)
APTO-001/Qwen3.5-9B-SafetyTuned (GGUF)… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/ja-safety-sft-dataset.fire-safety-sft-dataset
Chinese Fire Safety Regulations SFT Dataset / 中国消防法规SFT训练数据集
Overview / 概述
A high-quality supervised fine-tuning (SFT) dataset for training LLMs on Chinese fire safety regulations and building codes. Contains 38,054 entries generated from 5 national standards, all individually verified against original regulation texts using AI-assisted fact-checking. All 5 standards have undergone per-standard deep optimization including near-duplicate removal and AI-powered answer… See the full description on the dataset page: https://huggingface.co/datasets/sdzjoy/fire-safety-sft-dataset.agent-safety-bench
Agent Safety Bench (ASB)
ASB is a benchmark for evaluating the safety of tool-using LLM agents. Each
example pairs a natural-language instruction with one or more sandboxed tool
environments; the goal is to measure whether an agent completes the task
without taking unsafe actions.
This repository hosts the task data for ASB. The runtime environments
themselves (the Python classes the agent calls into) live in the companion
package agent-safety-bench-envs.
It ships two configs:… See the full description on the dataset page: https://huggingface.co/datasets/aradhye/agent-safety-bench.nemotron-nano2-safety-distill-gptoss
Nemotron Nano 2 Safety Distill — GPT-OSS
A distilled safety dataset produced using the Nemotron Nano 2 recipe with GPT-OSS-20B and GPT-OSS-120B as teacher models.
⚠️ Content Warning: This dataset includes potentially harmful prompts. Use responsibly for research purposes only.
Overview
This safety-focused distilled dataset was created by following the Nemotron Nano 2 safety recipe, adapted to use GPT-OSS-20B and GPT-OSS-120B as teacher models. Due to resource limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ericwang/nemotron-nano2-safety-distill-gptoss.bio-safety-peft-lora
CBRN Safety Alignment & PEFT-LoRA Fine-Tuning Dataset
This repository contains the synthetic instruction-tuning dataset (.jsonl) designed for parameter-efficient fine-tuning (PEFT-LoRA) of edge language models (specifically Qwen/Qwen2.5-1.5B-Instruct).
The dataset is curated to evaluate and modify model logit distributions, persona attributions, and dual-use safety boundaries regarding Chemical, Biological, Radiological, and Nuclear (CBRN) risk scenarios.
🤖 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/devsgnr/bio-safety-peft-lora.indic-safety-eval
IndicSafetyBench
Multi-turn safety evaluation benchmark for Indian languages. Tests models against 30 jailbreak techniques across 19 India-specific harm domains in 23 languages with 189 dialect varieties.
Stats
3,374 benchmark items
30 jailbreak techniques across 8 families
19 India harm domains (territorial, caste, religious, political, gender, etc.)
14 global harm categories (aligned with AILuminate v1.0 / Llama Guard 4)
23 languages (22 Indic + English)
~29%… See the full description on the dataset page: https://huggingface.co/datasets/anna-sarvam/indic-safety-eval.inconvenience-public-safety
inconvenience-public-safety
Three Korean public-safety registers converted to Korean braille under the 2017
revised rules (문화체육관광부 고시 제2017-15호). Every register is enumerated in
full, not sampled.
The registers are here because their documents are shaped differently, not
because three is more than one. A pesticide row is a filled-in form; a patient
leaflet is prose; an accident case is a paragraph an investigator wrote. Median
record length spans more than an order of magnitude… See the full description on the dataset page: https://huggingface.co/datasets/Yuyongkim/inconvenience-public-safety.multilingual-safety
Multilingual Safety Instructions
A multilingual extension of the safety-only instruction–refusal pairs released with the Safety-Tuned LLaMAs project. The original 1,000 harmful-prompt / refusal-response pairs (English) were translated into 11 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config.
Field
Description
prompt
Harmful user instruction (translated; en is the original).
output
Safe… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-safety.OTel-Safety
OTel-Safety
Dataset Summary
OTel-Safety is a specialized dataset for training large language models to abstain from answering when the retrieved context in a RAG pipeline is insufficient or irrelevant. It is part of the Open Telco (OTel) AI project, the largest open-source AI initiative in telecommunications, curated by over 100 domain experts from industry and academia.
In deployed RAG systems, a common failure mode is hallucination when the retrieval step returns… See the full description on the dataset page: https://huggingface.co/datasets/farbodtavakkoli/OTel-Safety.meddies-patient-safety
Meddies Patient Safety
A Vietnamese clinical red-team set: 22,336 synthetic patient queries probing five unsafe response modes, paired with 19,085 doctor-LLM responses that passed LLM-as-judge quality criteria.
[!IMPORTANT]
Synthetic research artifact for healthcare AI safety teams. Not medical advice. Not a substitute for IRB-grade clinical evaluation. Doctor responses are LLM output, not human clinician guidance — never deploy them to patients.
If you want to… See the full description on the dataset page: https://huggingface.co/datasets/Meddies/meddies-patient-safety.sp-sft-safety-180k
model-raising-pbsft-safety-180k
A constitution-aware paired SFT dataset of 182,688 safety-relevant prompts. Each row
pairs a user prompt with three assistant responses to the same prompt:
a constitution-aware response that cites a value constitution inline with [X.Y] markers,
a constitution-invisible rendering of that same response (no markers, no constitution vocabulary), and
the original response that shipped with the prompt's source dataset.
It is part of the Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/sp-sft-safety-180k.dailyconversationsThis dataset is synthetically generated using ChatGPT 3.5 to contain two-person multi-turn daily conversations with a various of topics (e.g.
travel, food, music, movie/TV, education, hobbies, family, sports, technology, books, etc.) Originally, this dataset is used to train
QuicktypeGPT, which is a GPT model to assist auto complete conversations.
Here is the full list of topics the conversation may cover.
safety_SFT_dataset_14k
Safety-SFT-Dataset-14K
1. Dataset Summary
Safety-SFT-14K is a professional-grade corpus of 14,000 instruction-tuning pairs specifically engineered for the Supervised Fine-Tuning (SFT) phase of Large Language Model (LLM) alignment. This dataset is optimized to train models to recognize, categorize, and appropriately refuse harmful requests across a diverse spectrum of safety violations while maintaining a helpful, neutral tone for benign queries.… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/safety_SFT_dataset_14k.
