guardrail
budgeted-guardrails-mlp-channel-interventions
MLP Channel Guardrails
This repository accompanies the Zenodo preprint
10.5281/zenodo.21003379,
Budgeted Guardrails for MLP Channel Interventions in Small Transformer
Language Models.
The study evaluates training-time MLP-channel interventions in small causal
Transformer language models. The supported claim is narrow: in the tested
small-model settings, low-amplitude MLP channel gates with strict event budgets
reduce trajectory damage relative to fixed event gating and dropout… See the full description on the dataset page: https://huggingface.co/datasets/cuber12/budgeted-guardrails-mlp-channel-interventions.detect-jailbreak
Content Warning:
This dataset contains unsafe model responses and user queries. Viewers may find the content disturbing.
Overview
Our evaluation dataset combines three existing datasets with custom augmentations to create a robust framework for assessing LLM vulnerabilities and defense effectiveness. The core components are the Verazuo dataset, the ZHX123 benchmark, and the Weapons of Mass Destruction Proxy (WMDP) dataset.
Credits and Citations
Our greatest… See the full description on the dataset page: https://huggingface.co/datasets/GuardrailsAI/detect-jailbreak.guardrail-tr
Guardrail-TR
Guardrail-TR (ytu-ce-cosmos/guardrail-tr) is a large-scale Turkish prompt-safety dataset for training and evaluating content-moderation / guardrail classifiers.
Dataset contains ~405K single-turn user prompts with:
a binary safe / unsafe label, and
multi-label hazard categories (a prompt may carry more than one category).
All released prompt text is Turkish. Rows that originated from English sources were processed, translated and, where relevant, culturally… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/guardrail-tr.arabic-guardrail
Arabic Guardrail — 250,842 rows, 12 classes
Defensive dataset for training Arabic prompt-safety classifiers. Each row is an incoming user
message and which of 12 safety classes it belongs to.
بالعربية: مجموعة بيانات عربية لتدريب نماذج تصنّف الرسائل الواردة قبل وصولها للمساعد الذكي.
Arabic guardrails were a gap. Hugging Face searches for Arabic jailbreak / safety /
prompt-injection datasets return zero results, and the one Arabic guardrail model… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-guardrail.guardrail-hard-negatives
Guardrail Hard Negatives (EN/TR)
A false-positive stress test for guardrails. A curated, paired benign/attack dataset for evaluating and training prompt-injection detectors and LLM guardrails. A bilingual false-positive challenge set: benign prompts that look like attacks (security researchers asking about injection, authorized admin actions, quoted payloads, legitimate roleplay) paired against real attacks, so you can measure the false-positive rate your users will actually… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/guardrail-hard-negatives.context-adherence-guardrail-10k
Context-Adherence Guardrail — training data (10,710)
Training data for a single-token RAG-groundedness guardrail. Each item is a
(question, context, response) triple with a human- or construction-derived
PASS/FAIL label under one Behavior Spec:
FAIL iff the response makes at least one factual claim unsupported by or
contradicting the retrieved context — truth in the real world is irrelevant
(strict grounding). PASS otherwise, including responses that decline to
answer for lack… See the full description on the dataset page: https://huggingface.co/datasets/aaryand/context-adherence-guardrail-10k.
