datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
libero_safetyAegis-AI-Content-Safety-Dataset-2.0
🛡️ Nemotron Content Safety Dataset V2
The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1.
To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.AgentHarm
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Maksym Andriushchenko1,†,*, Alexandra Souly2,*
Mateusz Dziemian1, Derek Duenas1, Maxwell Lin1, Justin Wang1, Dan Hendrycks1,§, Andy Zou1,¶,§, Zico Kolter1,¶, Matt Fredrikson1,¶,*
Eric Winsor2, Jerome Wynne2, Yarin Gal2,♯, Xander Davies2,♯,*
1Gray Swan AI, 2UK AI Safety Institute, *Core Contributor
†EPFL, §Center for AI Safety, ¶Carnegie Mellon University, ♯University of Oxford
Paper: https://arxiv.org/abs/2410.09024… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/AgentHarm.Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.indoor-safety-hazard-detection-and-work-zone-monitoring
Indoor Safety Hazard Detection & Work-Zone Monitoring
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/indoor-safety-hazard-detection-and-work-zone-monitoring.Aegis-AI-Content-Safety-Dataset-1.0
🛡️ Nemotron Content Safety Dataset V1
Nemotron Content Safety Dataset V1, formerly known as Aegis AI Content Safety Dataset, is an open-source content safety dataset (CC-BY-4.0), which adheres to Nvidia's content safety taxonomy, covering 13 critical risk categories (see Dataset Description).
Dataset Details
Dataset Description
Nemotron Content Safety Dataset V1 is comprised of approximately 11,000 manually annotated interactions between humans and LLMs, split… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-1.0.lie-detection-rollouts
Lie Detection Rollouts
Assistant completions across many open-weight models on the lie-detection
evaluation suite used by the
deception research pipeline. One subset per model,
one split per task.
Columns
messages — list of OpenAI-style messages. Each message has:
role: system | user | assistant
content: final message text
reasoning_content: chain-of-thought for reasoning models, None otherwise
is_lie — ground-truth label from the is_deceptive scorer:
lie |… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/lie-detection-rollouts.gridstar-safety-dataSafety-Toxicity-Detection
SEA Toxicity Detection
SEA Toxicity Detection evaluates a model's ability to identify toxic content such as hate speech and abusive language in text. It is sampled from MLHSD for Indonesian, TTD for Thai, and ViHSD for Vietnamese.
Supported Tasks and Leaderboards
SEA Toxicity Detection is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Thai… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Safety-Toxicity-Detection.Safety-helmet-datasetlatent-mas-safety-dataset-seq-qwen3-4b
LatentMAS Safety Dataset — Phase 0
Latent states, model completions, and safety labels from a Qwen3-4B
latent-MAS pipeline (Planner → Critic → Refiner → Judger, inter-agent
messages passed as hidden-state vectors) evaluated on prompts from
public safety benchmarks. Intended for training a latent safety value
model and for probing / interpretability work on multi-agent latent
reasoning.
What's in it
195,589 rollouts from 14 prompt sources, greedy decode… See the full description on the dataset page: https://huggingface.co/datasets/asatheesh/latent-mas-safety-dataset-seq-qwen3-4b.MM-SafetyBenchWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements.
Usage and License Notices: The dataset is intended and licensed for research use only. They are also restricted to uses that follow the license agreement GPT-4 and Stable Diffusion. The dataset is CC BY NC 4.0 (allowing only non-commercial use).
Data Source: For more information about the dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/MM-SafetyBench.Safety-helmet-datasetNemotron-Safety-Guard-Dataset-v3
Dataset Description:
The Nemotron-Safety-Guard-Dataset-v3 (formerly known as Nemotron-Content-Safety-Dataset-Multilingual-v1) is a large, high-quality safety dataset designed for training multilingual LLM safety guard models. It comprises approximately 514,617 samples across 12 languages: English, Arabic, German, Spanish, French, Hindi, Japanese, Thai, Mandarin, Dutch, Italian, and Korean.
This dataset is primarily synthetically generated using the CultureGuard pipeline, which… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Safety-Guard-Dataset-v3.MMA-SafetyBenchsafety-classifications
Safety Annotations for dolma3_mix
Safety score annotations for a 20K-shard subset of allenai/dolma3_mix-6T using
locuslab/safety-classifier_gte-large-en-v1.5.
Schema
Column
Type
Description
id
string
Row identifier (matches source dataset)
safety_score
int8
Argmax safety class (0-5)
safety_probs
list[float32]
Full 6-class probability distribution
Safety scale
Score
Label
Count
Percentage
0
safe
302,972,734
77.39%
1… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/safety-classifications.SafetyBenchSafetyBench is a comprehensive benchmark for evaluating the safety of LLMs, which comprises 11,435 diverse multiple choice questions spanning across 7 distinct categories of safety concerns. Notably, SafetyBench also incorporates both Chinese and English data, facilitating the evaluation in both languages.
Please visit our GitHub and website or check our paper for more details.
We release three differents test sets including Chinese testset (test_zh.json), English testset (test_en.json) and… See the full description on the dataset page: https://huggingface.co/datasets/thu-coai/SafetyBench.prompt-injection-safetylatent-mas-safety-dataset-seq-qwen3-14bMM-SafetyBench-plus-plus
MM-SafetyBench++
Project Page | Paper | Code
MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent.
Dataset Summary
For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.Nemotron-Content-Safety-Audio-Dataset
Nemotron Content Safety Audio Dataset
Dataset Description
The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories.
LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.construction-safety-object-detection
Dataset Labels
['barricade', 'dumpster', 'excavators', 'gloves', 'hardhat', 'mask', 'no-hardhat', 'no-mask', 'no-safety vest', 'person', 'safety net', 'safety shoes', 'safety vest', 'dump truck', 'mini-van', 'truck', 'wheel loader']
Number of Images
{'train': 307, 'valid': 57, 'test': 34}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/keremberke/construction-safety-object-detection.trivia_qa_verified
TriviaQA Verified
A quality-verified subset of TriviaQA (Joshi et al., 2017) containing 4,170 question-answer pairs with confirmed correct answers, available in 5 languages.
Splits
Split
Language
Rows
english
English
4,170
mandarin
Mandarin Chinese
4,170
japanese
Japanese
4,170
arabic
Arabic
4,170
french
French
4,170
validation
English
3,381
The validation split contains a separate set of verified English questions (no overlap with other splits)… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/trivia_qa_verified.safety-data
⚠️ Content Warning: This dataset contains sensitive prompts and model
responses that include harmful, offensive, and dangerous content. It is intended
for safety research.
Dataset Card for Safety-IRT
Data for "Why Do Safety Guardrails Degrade Across Languages?"
Contains 1.9M graded responses from 61 model configurations across 10 languages,
along with anchor selections, judge validation data, and native speaker translation ratings.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/safety-irt/safety-data.Nemotron-SFT-Safety-v2
Dataset Description:
The Nemotron-SFT-Safety-v2 data is designed to align models to be robust against a variety of safety and security concerns that may arise in unaligned large language models.This dataset is a collection of:
A hybrid (open-source and synthetically generated) collection of prompts designed to elicit different model vulnerabilities, and
Synthetically generated responses designed to steer model behavior towards safety-aligned values and enhance model robustness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Safety-v2.construction-safety-gsnvb
Construction Safety Gsnvb
This dataset is part of the Roboflow 100 benchmark, a diverse collection of 100 object detection datasets spanning 7 imagery domains.
Dataset Statistics
Split
Images
Train
997
Validation
119
Test
90
Total
1,206
Classes (5)
helmet
no-helmet
no-vest
person
vest
Usage
With LibreYOLO
from libreyolo import LIBREYOLO
# Load a model
model = LIBREYOLO(model_path="libreyoloXnano.pt")
# Train on… See the full description on the dataset page: https://huggingface.co/datasets/LibreYOLO/construction-safety-gsnvb.Nemotron-3.5-Content-Safety-Dataset
Nemotron 3.5 Content Safety Dataset
Dataset Description:
Nemotron 3.5 Content Safety Dataset is a hybrid real/synthetic supervised instruction dataset for content-safety classification of human and assistant interactions. The dataset contains text-only and image-grounded single-turn conversations. Each example asks a classifier to determine user safety, response safety, and harmful categories; a subset also covers topic-following classification. Some training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-3.5-Content-Safety-Dataset.XL-SafetyBench
XL-SafetyBench
A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
⚠️ Content Warning: This dataset contains adversarial prompts and
culturally sensitive content for safety and cultural-evaluation research.
By using this dataset, you agree to use it solely for research purposes
and not for malicious applications.
Paper: https://arxiv.org/abs/2605.05662
Eval Code: github.com/AIM-Intelligence/XL-SafetyBench
Overview… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/XL-SafetyBench.libero_safety_assetsSafety-Prompts
Dataset Card for Dataset Name
GitHub Repository: https://github.com/thu-coai/Safety-Prompts
Paper: https://arxiv.org/abs/2304.10436
