guardrails
detect-jailbreak
Content Warning:
This dataset contains unsafe model responses and user queries. Viewers may find the content disturbing.
Overview
Our evaluation dataset combines three existing datasets with custom augmentations to create a robust framework for assessing LLM vulnerabilities and defense effectiveness. The core components are the Verazuo dataset, the ZHX123 benchmark, and the Weapons of Mass Destruction Proxy (WMDP) dataset.
Credits and Citations
Our greatest… See the full description on the dataset page: https://huggingface.co/datasets/GuardrailsAI/detect-jailbreak.budgeted-guardrails-mlp-channel-interventions
MLP Channel Guardrails
This repository accompanies the Zenodo preprint
10.5281/zenodo.21003379,
Budgeted Guardrails for MLP Channel Interventions in Small Transformer
Language Models.
The study evaluates training-time MLP-channel interventions in small causal
Transformer language models. The supported claim is narrow: in the tested
small-model settings, low-amplitude MLP channel gates with strict event budgets
reduce trajectory damage relative to fixed event gating and dropout… See the full description on the dataset page: https://huggingface.co/datasets/cuber12/budgeted-guardrails-mlp-channel-interventions.content-moderation
Note:
This dataset contains the EVAL portion of the Jigsaw Toxic Comment Dataset.
It should be used for model evaluation. For training, one can use the original Jigsaw dataset: https://huggingface.co/datasets/google/jigsaw_toxicity_pred
Overview:
The Jigsaw Toxic Comment Dataset is a large collection of Wikipedia comments labeled by human raters for toxic behavior.
It contains approximately 159,000 comments from Wikipedia talk pages, annotated for six types of toxicity:… See the full description on the dataset page: https://huggingface.co/datasets/GuardrailsAI/content-moderation.Evaluation-Dataset-of-AI-Agent-Security-Guardrails
DKnownAI Agent Security Evaluation Dataset
Data Fields
Field
Type
Description
text
string
The adversarial input (prompt) to be evaluated by a security guardrail
action
string
Human-annotated label: blocked or allowed
Citation
@misc{li2026comparativeevaluationaiagent,
title={A Comparative Evaluation of AI Agent Security Guardrails},
author={Qi Li and Jiu Li and Pingtao Wei and Jianjun Xu and Xueyi Wei and Jiwei Shi and Xuan… See the full description on the dataset page: https://huggingface.co/datasets/CaiZhiTech/Evaluation-Dataset-of-AI-Agent-Security-Guardrails.agent-symbolic-guardrailsThis dataset contains data associated with the paper Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility.
Code: https://github.com/hyn0027/agent-symbolic-guardrails
Subsets
literature_review
This subset contains the metadata of the systematic literature review data. Details are discussed in Section 3 in the paper.
adversarial_MedAgentBench
This subset contains the adversarial tasks we… See the full description on the dataset page: https://huggingface.co/datasets/hyn0027D/agent-symbolic-guardrails.guardrails-dataset
