CoolFace
20 results

guardrails

GuardrailsAI /detect-jailbreak Content Warning: This dataset contains unsafe model responses and user queries. Viewers may find the content disturbing. Overview Our evaluation dataset combines three existing datasets with custom augmentations to create a robust framework for assessing LLM vulnerabilities and defense effectiveness. The core components are the Verazuo dataset, the ZHX123 benchmark, and the Weapons of Mass Destruction Proxy (WMDP) dataset. Credits and Citations Our greatest… See the full description on the dataset page: https://huggingface.co/datasets/GuardrailsAI/detect-jailbreak.texttext-classification10K<n<100K7 likes279 downloads2y agoHugging Facecuber12 /budgeted-guardrails-mlp-channel-interventions MLP Channel Guardrails This repository accompanies the Zenodo preprint 10.5281/zenodo.21003379, Budgeted Guardrails for MLP Channel Interventions in Small Transformer Language Models. The study evaluates training-time MLP-channel interventions in small causal Transformer language models. The supported claim is narrow: in the tested small-model settings, low-amplitude MLP channel gates with strict event budgets reduce trajectory damage relative to fixed event gating and dropout… See the full description on the dataset page: https://huggingface.co/datasets/cuber12/budgeted-guardrails-mlp-channel-interventions.0 likes235 downloads3mo agoHugging FaceGuardrailsAI /content-moderation Note: This dataset contains the EVAL portion of the Jigsaw Toxic Comment Dataset. It should be used for model evaluation. For training, one can use the original Jigsaw dataset: https://huggingface.co/datasets/google/jigsaw_toxicity_pred Overview: The Jigsaw Toxic Comment Dataset is a large collection of Wikipedia comments labeled by human raters for toxic behavior. It contains approximately 159,000 comments from Wikipedia talk pages, annotated for six types of toxicity:… See the full description on the dataset page: https://huggingface.co/datasets/GuardrailsAI/content-moderation.text-classification1 likes100 downloads2y agoHugging FaceCaiZhiTech /Evaluation-Dataset-of-AI-Agent-Security-Guardrails DKnownAI Agent Security Evaluation Dataset Data Fields Field Type Description text string The adversarial input (prompt) to be evaluated by a security guardrail action string Human-annotated label: blocked or allowed Citation @misc{li2026comparativeevaluationaiagent, title={A Comparative Evaluation of AI Agent Security Guardrails}, author={Qi Li and Jiu Li and Pingtao Wei and Jianjun Xu and Xueyi Wei and Jiwei Shi and Xuan… See the full description on the dataset page: https://huggingface.co/datasets/CaiZhiTech/Evaluation-Dataset-of-AI-Agent-Security-Guardrails.texttext-classification1K<n<10K1 likes98 downloads5mo agoHugging Facehyn0027D /agent-symbolic-guardrailsThis dataset contains data associated with the paper Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility. Code: https://github.com/hyn0027/agent-symbolic-guardrails Subsets literature_review This subset contains the metadata of the systematic literature review data. Details are discussed in Section 3 in the paper. adversarial_MedAgentBench This subset contains the adversarial tasks we… See the full description on the dataset page: https://huggingface.co/datasets/hyn0027D/agent-symbolic-guardrails.textothern<1K1 likes76 downloads3mo agoHugging Facehuyhoangdinhcong /guardrails-datasettabular100K<n<1M0 likes70 downloads11mo agoHugging Face