LLM-safety
HEx-PHI
HEx-PHI: Human-Extended Policy-Oriented Harmful Instruction Benchmark
This dataset contains 330 harmful instructions (30 examples x 11 prohibited categories) for LLM harmfulness evaluation.
In our work "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!", to comprehensively cover as many harmfulness categories as possible,
we develop this new safety evaluation benchmark directly based on the exhaustive lists of prohibited use cases found in… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Tuning-Safety/HEx-PHI.MaliciousInstructllm-sfm-safety-eval
LLM x SFM Safety Evaluation
When a general-purpose language model interprets the output of a specialist
science foundation model (a protein, genomic, RNA, or chemistry model), does its
safety behavior recognize the scientific content, or only the surface form of the
request?
This repository is the empirical core of a study of that question: the evaluation
harness, the redacted aggregate results, and the measurement specifications behind
four findings about how deployed Claude… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/llm-sfm-safety-eval.llm-safety-instruct-v7mlcommons-ai-safety-synth
MLCommons AI Safety Synthesized Dataset
Synthesized training data for AI safety classifiers based on the MLCommons AI Safety Hazard Taxonomy.
Dataset Description
This dataset contains 12,000 synthesized unsafe prompts across 6 hazard categories, designed to augment training data for content safety classifiers. Each category contains 2,000 balanced samples.
Hazard Categories (MLCommons AI Safety Taxonomy)
Category
Description
Samples… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/mlcommons-ai-safety-synth.bayesian-llm-safety-inference
Bayesian Latent Safety-Trait Dataset
Summary
This dataset supports Bayesian latent-trait analysis of language-model safety behavior.
It contains 90 benchmark-derived roots, three matched prompt variants per root, responses
from four target models over five runs, two independent LLM ratings per response, and one
human rating for a stratified 540-response calibration subset.
The three dimensions are harmful compliance, sycophancy, and agentic protocol violation.… See the full description on the dataset page: https://huggingface.co/datasets/Charly-X/bayesian-llm-safety-inference.
