datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm-sfm-safety-eval
LLM x SFM Safety Evaluation
When a general-purpose language model interprets the output of a specialist
science foundation model (a protein, genomic, RNA, or chemistry model), does its
safety behavior recognize the scientific content, or only the surface form of the
request?
This repository is the empirical core of a study of that question: the evaluation
harness, the redacted aggregate results, and the measurement specifications behind
four findings about how deployed Claude… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/llm-sfm-safety-eval.mlcommons-ai-safety-synth
MLCommons AI Safety Synthesized Dataset
Synthesized training data for AI safety classifiers based on the MLCommons AI Safety Hazard Taxonomy.
Dataset Description
This dataset contains 12,000 synthesized unsafe prompts across 6 hazard categories, designed to augment training data for content safety classifiers. Each category contains 2,000 balanced samples.
Hazard Categories (MLCommons AI Safety Taxonomy)
Category
Description
Samples… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/mlcommons-ai-safety-synth.llm_physical_safety_benchmark
LLM Physical Safety Benchmark in Drone Control
This benchmark consists of four datasets designed to evaluate the performance of Large Language Models (LLMs) in controlling drones and their vulnerability to physical attacks. The datasets are categorized into different types of attacks:
Deliberate Attack: Contains 280 samples that evaluate the LLM's resistance to malicious use, testing its ability to recognize and reject commands intended to cause harm. Subcategories include Direct… See the full description on the dataset page: https://huggingface.co/datasets/TrustSafeAI/llm_physical_safety_benchmark.llm_physical_safety_benchmark
LLM Physical Safety Benchmark in Drone Control
This benchmark consists of four datasets designed to evaluate the performance of Large Language Models (LLMs) in controlling drones and their vulnerability to physical attacks. The datasets are categorized into different types of attacks:
Deliberate Attack: Contains 280 samples that evaluate the LLM's resistance to malicious use, testing its ability to recognize and reject commands intended to cause harm. Subcategories include Direct… See the full description on the dataset page: https://huggingface.co/datasets/kumitang/llm_physical_safety_benchmark.amalia-Nemotron-SFT-Safety-v1
AMALIA Nemotron-SFT-Safety-v1
Version of the nvidia/Nemotron-SFT-Safety-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage.
This dataset went through a processing pipeline to remove entries that reference other LLMs or research labs;
Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Safety-v1
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SFT-Safety-v1.repro-a-coin-flip-for-safety-llm-judges-fail-to-reliably-measure-adversarial-robustness
Reproduction: A Coin Flip for Safety - LLM Judges Fail to Reliably Measure Adversarial Robustness
Paper Information
Title: A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
OpenReview ID: RgnoWsmYBM
Conference: ICML 2026
Task: Evaluate 4 LLM-based safety judges on 6,642 human-verified adversarial prompts
Reproduction Summary
This paper's all 6 claims require LLM inference on proprietary adversarial prompts and… See the full description on the dataset page: https://huggingface.co/datasets/sabaridsnfuji/repro-a-coin-flip-for-safety-llm-judges-fail-to-reliably-measure-adversarial-robustness.
