CoolFace
Datasetpublic

BushNate/llm-refusal-evaluation

🛡️ LLM Refusal Evaluation Benchmark This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite. The prompts are organized into three groups: Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse. Chinese Sensitive Topics — prompts that may be censored by China-aligned models. Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse. 📌 Contents Safety Benchmarks JailbreakBench SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/BushNate/llm-refusal-evaluation.

sourceHugging Faceupdated 2mo agoView on Hugging Face
1likes85downloads
1 commits on main
73ea5222mo ago

Duplicate from MultiverseComputingCAI/llm-refusal-evaluation

BushNate, Iker