CoolFace
20 results

guardrail

cuber12 /budgeted-guardrails-mlp-channel-interventions MLP Channel Guardrails This repository accompanies the Zenodo preprint 10.5281/zenodo.21003379, Budgeted Guardrails for MLP Channel Interventions in Small Transformer Language Models. The study evaluates training-time MLP-channel interventions in small causal Transformer language models. The supported claim is narrow: in the tested small-model settings, low-amplitude MLP channel gates with strict event budgets reduce trajectory damage relative to fixed event gating and dropout… See the full description on the dataset page: https://huggingface.co/datasets/cuber12/budgeted-guardrails-mlp-channel-interventions.0 likes232 downloads3mo agoHugging FaceGuardrailsAI /detect-jailbreak Content Warning: This dataset contains unsafe model responses and user queries. Viewers may find the content disturbing. Overview Our evaluation dataset combines three existing datasets with custom augmentations to create a robust framework for assessing LLM vulnerabilities and defense effectiveness. The core components are the Verazuo dataset, the ZHX123 benchmark, and the Weapons of Mass Destruction Proxy (WMDP) dataset. Credits and Citations Our greatest… See the full description on the dataset page: https://huggingface.co/datasets/GuardrailsAI/detect-jailbreak.texttext-classification10K<n<100K7 likes227 downloads2y agoHugging Faceytu-ce-cosmos /guardrail-tr Guardrail-TR Guardrail-TR (ytu-ce-cosmos/guardrail-tr) is a large-scale Turkish prompt-safety dataset for training and evaluating content-moderation / guardrail classifiers. Dataset contains ~405K single-turn user prompts with: a binary safe / unsafe label, and multi-label hazard categories (a prompt may carry more than one category). All released prompt text is Turkish. Rows that originated from English sources were processed, translated and, where relevant, culturally… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/guardrail-tr.texttext-classification100K<n<1M7 likes207 downloads2mo agoHugging Faceoddadmix /arabic-guardrail Arabic Guardrail — 250,842 rows, 12 classes Defensive dataset for training Arabic prompt-safety classifiers. Each row is an incoming user message and which of 12 safety classes it belongs to. بالعربية: مجموعة بيانات عربية لتدريب نماذج تصنّف الرسائل الواردة قبل وصولها للمساعد الذكي. Arabic guardrails were a gap. Hugging Face searches for Arabic jailbreak / safety / prompt-injection datasets return zero results, and the one Arabic guardrail model… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-guardrail.tabulartext-classification100K<n<1M0 likes180 downloads23d agoHugging Face3nesdeniz /guardrail-hard-negatives Guardrail Hard Negatives (EN/TR) A false-positive stress test for guardrails. A curated, paired benign/attack dataset for evaluating and training prompt-injection detectors and LLM guardrails. A bilingual false-positive challenge set: benign prompts that look like attacks (security researchers asking about injection, authorized admin actions, quoted payloads, legitimate roleplay) paired against real attacks, so you can measure the false-positive rate your users will actually… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/guardrail-hard-negatives.tabulartext-classification1K<n<10K3 likes136 downloads1mo agoHugging Faceaaryand /context-adherence-guardrail-10k Context-Adherence Guardrail — training data (10,710) Training data for a single-token RAG-groundedness guardrail. Each item is a (question, context, response) triple with a human- or construction-derived PASS/FAIL label under one Behavior Spec: FAIL iff the response makes at least one factual claim unsupported by or contradicting the retrieved context — truth in the real world is irrelevant (strict grounding). PASS otherwise, including responses that decline to answer for lack… See the full description on the dataset page: https://huggingface.co/datasets/aaryand/context-adherence-guardrail-10k.texttext-classification10K<n<100K1 likes111 downloads1mo agoHugging Face