CoolFace
20 results

LLM-safety

LLM-Tuning-Safety /HEx-PHIgated HEx-PHI: Human-Extended Policy-Oriented Harmful Instruction Benchmark This dataset contains 330 harmful instructions (30 examples x 11 prohibited categories) for LLM harmfulness evaluation. In our work "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!", to comprehensively cover as many harmfulness categories as possible, we develop this new safety evaluation benchmark directly based on the exhaustive lists of prohibited use cases found in… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Tuning-Safety/HEx-PHI.text-generationn<1K66 likes312 downloads2y agoHugging FaceLLMSafety /MaliciousInstructtextn<1K0 likes257 downloads7mo agoHugging Facejang1563 /llm-sfm-safety-eval LLM x SFM Safety Evaluation When a general-purpose language model interprets the output of a specialist science foundation model (a protein, genomic, RNA, or chemistry model), does its safety behavior recognize the scientific content, or only the surface form of the request? This repository is the empirical core of a study of that question: the evaluation harness, the redacted aggregate results, and the measurement specifications behind four findings about how deployed Claude… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/llm-sfm-safety-eval.texttext-classification10K<n<100K0 likes202 downloads13d agoHugging FaceUranus /llm-safety-instruct-v7text100K<n<1M0 likes133 downloads2y agoHugging Facellm-semantic-router /mlcommons-ai-safety-synth MLCommons AI Safety Synthesized Dataset Synthesized training data for AI safety classifiers based on the MLCommons AI Safety Hazard Taxonomy. Dataset Description This dataset contains 12,000 synthesized unsafe prompts across 6 hazard categories, designed to augment training data for content safety classifiers. Each category contains 2,000 balanced samples. Hazard Categories (MLCommons AI Safety Taxonomy) Category Description Samples… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/mlcommons-ai-safety-synth.texttext-classification10K<n<100K1 likes104 downloads8mo agoHugging FaceCharly-X /bayesian-llm-safety-inference Bayesian Latent Safety-Trait Dataset Summary This dataset supports Bayesian latent-trait analysis of language-model safety behavior. It contains 90 benchmark-derived roots, three matched prompt variants per root, responses from four target models over five runs, two independent LLM ratings per response, and one human rating for a stratified 540-response calibration subset. The three dimensions are harmful compliance, sycophancy, and agentic protocol violation.… See the full description on the dataset page: https://huggingface.co/datasets/Charly-X/bayesian-llm-safety-inference.tabulartext-classification10K<n<100K0 likes76 downloads2mo agoHugging Face