APTO-001/APTO-SorryBench-JA
APTO-SorryBench-JA APTO-SorryBench-JA is a Japanese translated and annotated version of the original Sorry-Bench dataset for AI safety evaluation research. Sorry-Bench is a benchmark designed to evaluate whether Large Language Models (LLMs) appropriately refuse harmful requests while still providing helpful responses to safe requests. The original English prompts are preserved alongside the Japanese translations to improve traceability and facilitate comparison with the original… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/APTO-SorryBench-JA.
012
No card is published for this repository, or it could not be fetched from Hugging Face right now.
