datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MaliciousInstructllm-sfm-safety-eval
LLM x SFM Safety Evaluation
When a general-purpose language model interprets the output of a specialist
science foundation model (a protein, genomic, RNA, or chemistry model), does its
safety behavior recognize the scientific content, or only the surface form of the
request?
This repository is the empirical core of a study of that question: the evaluation
harness, the redacted aggregate results, and the measurement specifications behind
four findings about how deployed Claude… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/llm-sfm-safety-eval.llm-safety-instruct-v7mlcommons-ai-safety-synth
MLCommons AI Safety Synthesized Dataset
Synthesized training data for AI safety classifiers based on the MLCommons AI Safety Hazard Taxonomy.
Dataset Description
This dataset contains 12,000 synthesized unsafe prompts across 6 hazard categories, designed to augment training data for content safety classifiers. Each category contains 2,000 balanced samples.
Hazard Categories (MLCommons AI Safety Taxonomy)
Category
Description
Samples… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/mlcommons-ai-safety-synth.bayesian-llm-safety-inference
Bayesian Latent Safety-Trait Dataset
Summary
This dataset supports Bayesian latent-trait analysis of language-model safety behavior.
It contains 90 benchmark-derived roots, three matched prompt variants per root, responses
from four target models over five runs, two independent LLM ratings per response, and one
human rating for a stratified 540-response calibration subset.
The three dimensions are harmful compliance, sycophancy, and agentic protocol violation.… See the full description on the dataset page: https://huggingface.co/datasets/Charly-X/bayesian-llm-safety-inference.llm_physical_safety_benchmark
LLM Physical Safety Benchmark in Drone Control
This benchmark consists of four datasets designed to evaluate the performance of Large Language Models (LLMs) in controlling drones and their vulnerability to physical attacks. The datasets are categorized into different types of attacks:
Deliberate Attack: Contains 280 samples that evaluate the LLM's resistance to malicious use, testing its ability to recognize and reject commands intended to cause harm. Subcategories include Direct… See the full description on the dataset page: https://huggingface.co/datasets/TrustSafeAI/llm_physical_safety_benchmark.llm_physical_safety_benchmark
LLM Physical Safety Benchmark in Drone Control
This benchmark consists of four datasets designed to evaluate the performance of Large Language Models (LLMs) in controlling drones and their vulnerability to physical attacks. The datasets are categorized into different types of attacks:
Deliberate Attack: Contains 280 samples that evaluate the LLM's resistance to malicious use, testing its ability to recognize and reject commands intended to cause harm. Subcategories include Direct… See the full description on the dataset page: https://huggingface.co/datasets/kumitang/llm_physical_safety_benchmark.SorryBenchHarmBenchturkish-llm-authority-bypass-safety-sft
Turkish LLM Safety Dataset — Authority & System Command Bypass Refusal
Kod adı: TR-Auth-Bypass-Refusal-v1
Dil: Türkçe (tr)
Format: Hugging Face / Unsloth chat template uyumlu
🇹🇷 Türkçe Açıklama
Amaç
Bu veri seti, büyük dil modellerinin (LLM) güvenlik bariyerlerini (guardrails) aşmaya yönelik yetki süistimali ve sistem komutu bypass saldırılarını tespit edip güvenli biçimde reddetmesi için hazırlanmış bir Supervised Fine-Tuning (SFT) veri setidir.… See the full description on the dataset page: https://huggingface.co/datasets/sadecebirisii/turkish-llm-authority-bypass-safety-sft.XSTestpwc747_a10865__paper__P02__2023__high__llm_safety
Northwind Support Tickets Archive
A derived dataset combining service interaction logs with survey responses for support ticket analysis.
Upstream Sources
This dataset is derived from the following upstream source datasets:
Northwind Service Interaction Logs (TianfuXinqu/pwc747_a10865__paper__P05__2022__high__llm_safety)
Northwind Customer Survey Responses (TianfuXinqu/pwc747_a10865__paper__P06__2018__low__federated_learning)
Commercial Use… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/pwc747_a10865__paper__P02__2023__high__llm_safety.llm-safety-dpo-v3or-benchllm-safety-instruct-v6NuminaMath-CoT-100kOpenCodeInstruct-50kamalia-Nemotron-SFT-Safety-v1
AMALIA Nemotron-SFT-Safety-v1
Version of the nvidia/Nemotron-SFT-Safety-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage.
This dataset went through a processing pipeline to remove entries that reference other LLMs or research labs;
Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Safety-v1
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SFT-Safety-v1.hh-dpoArenaHardMT-Benchsorry-bench-rolellm-safety-japanese-multiturn-dataset株式会社APTOはLLMの安全性の性能を改善させるためのデータセットの開発を行いました。
近年、LLMの性能は飛躍的に向上しており、また、LLMの安全性は長らく注目を集めています。例えば、Anthropic社のClaudeでは、不正な入力に対しClaudeを保護する憲法AIが導入されているなど、安全性は世界的にも重要視されています。※1
しかし、今でも課題視されている点もあり、例えばGPT-5においても特定の条件で安全な会話ができなくなるケースも見られています。※2
日本国内でも例外なくLLMの安全性に対して課題感を持ちながらも安全性の向上への取り組みが見られます。
そのような国内のニーズに答えるべく、日本語で構成された安全性向上のためのデータセットを公開いたしました。
データの内容
データセットの構成として、ターン数がかさむほど安全性が劣化するという点に着目して、マルチターンのデータセットを作成しました。以下のサンプルのようなフォーマットのデータ構成となっております。
{
"question_turn1":… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/llm-safety-japanese-multiturn-dataset.AdvBenchrepro-a-coin-flip-for-safety-llm-judges-fail-to-reliably-measure-adversarial-robustness
Reproduction: A Coin Flip for Safety - LLM Judges Fail to Reliably Measure Adversarial Robustness
Paper Information
Title: A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
OpenReview ID: RgnoWsmYBM
Conference: ICML 2026
Task: Evaluate 4 LLM-based safety judges on 6,642 human-verified adversarial prompts
Reproduction Summary
This paper's all 6 claims require LLM inference on proprietary adversarial prompts and… See the full description on the dataset page: https://huggingface.co/datasets/sabaridsnfuji/repro-a-coin-flip-for-safety-llm-judges-fail-to-reliably-measure-adversarial-robustness.LLM-AI-Safety-Response-Classification
ASRCD — AI Safety Response Classification Dataset
Understanding AI Decision Making, Harm Detection, and Response Strategy in Real-World LLM Interactions
Author: Umair SaeedVersion: 1.0Total Rows: 1,000Format: CSVLanguage: EnglishTask Type: Multi-label Text ClassificationLicense: Research and Educational Use Only
About Dataset
ASRCD — AI Safety Response Classification Dataset
Overview
This dataset is designed to train and evaluate… See the full description on the dataset page: https://huggingface.co/datasets/umairpy/LLM-AI-Safety-Response-Classification.LLM_SAFETY_CHECK_OOD
