llm-guard
details_guardrail__llama-2-7b-guanaco-instruct-sharded
Dataset Card for Evaluation run of guardrail/llama-2-7b-guanaco-instruct-sharded
Dataset Summary
Dataset automatically created during the evaluation run of model guardrail/llama-2-7b-guanaco-instruct-sharded on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_guardrail__llama-2-7b-guanaco-instruct-sharded.llm-prompt-guard-tuning-corpus
llm-prompt-guard tuning corpus
Prompt-injection detection corpus used to tune the
llm-prompt-guard pattern
set. Two JSONL files:
attacks.jsonl — 198 rows, label: 1. Injection payloads grouped by
attack category (instruction override, role hijacking, jailbreak,
unicode/homoglyph/tag-block smuggling, encoding bypass, and more).
benign.jsonl — 1,310 rows, label: 0. Ordinary user input across
seven domains, including phrasing that superficially resembles an
attack ("ignore the… See the full description on the dataset page: https://huggingface.co/datasets/shanemhamilton/llm-prompt-guard-tuning-corpus.llm_guard_datasetCounseling-LLM-guardrailDormamu_GVD-1_Guardrail_Vulnerability_Dataset_for_LLM_Alignment
Dormamu GVD-1: Alignment Benchmark Dataset
Introduction
Dormamu is an advanced AI alignment framework developed to enhance the safety and ethical robustness of large language models (LLMs). It extends the OWASP LLM Top 10 vulnerabilities with proprietary categories, such as DormamuX1 for existential risk simulations and DormamuX2 for cultural alignment testing.
The Guardrail Vulnerability Dataset version 1 (GVD-1) is a comprehensive collection of test cases designed to benchmark LLM… See the full description on the dataset page: https://huggingface.co/datasets/Dormamu-Labs/Dormamu_GVD-1_Guardrail_Vulnerability_Dataset_for_LLM_Alignment.
