CoolFace
20 results

redteam

jash-ai /agentic-redteam-benchmark agentic-redteam-benchmark v0.8 preview · 2,288 multi-step agent trajectories · 513 hand-authored gold + 1,775 provenance-flagged augmented. A per-step benchmark that scores whether a verifier catches drift inside an agent's trajectory — not whether a prompt is harmful. 📦 Code, eval harness & issues: github.com/Alkur123/agentic-redteam-benchmark · 📄 Paper: A Per-Step Trajectory Benchmark for AI-Agent Governance Verifiers and a Corrected Catch-at-Drift Metric (Aegis AI, 2026)… See the full description on the dataset page: https://huggingface.co/datasets/jash-ai/agentic-redteam-benchmark.texttext-classification1K<n<10K2 likes1.3k downloads21d agoHugging FaceCohereLabs /aya_redteaming Dataset Card for Aya Red-teaming Dataset Details The Aya Red-teaming dataset is a human-annotated multilingual red-teaming dataset consisting of harmful prompts in 8 languages across 9 different categories of harm with explicit labels for "global" and "local" harm. Curated by: Professional compensated annotators Languages: Arabic, English, Filipino, French, Hindi, Russian, Serbian and Spanish License: Apache 2.0 Paper: arxiv link Harm Categories:… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_redteaming.text1K<n<10K36 likes1k downloads7mo agoHugging FaceMMInstruction /RedTeamingVLMRed Teaming Viusal Language Models17 likes611 downloads2y agoHugging Facewalledai /TDC23-RedTeaming TDC 2023 (LLM Edition) - Red Teaming Track This is the combined dev and test set from the Red Teaming Track of TDC 2023. Citation If find this dataset useful, please cite the following work: @inproceedings{tdc2023, title={TDC 2023 (LLM Edition): The Trojan Detection Challenge}, author={Mantas Mazeika and Andy Zou and Norman Mu and Long Phan and Zifan Wang and Chunru Yu and Adam Khoja and Fengqing Jiang and Aidan O'Gara and Ellie Sakhaee and Zhen Xiang and Arezoo… See the full description on the dataset page: https://huggingface.co/datasets/walledai/TDC23-RedTeaming.textn<1K8 likes391 downloads2y agoHugging FaceQwovadis /red-team-appsec-benchmark 🛡️ AI-SaaS AppSec Benchmark — v25 Открытый held-out бенчмарк для оценки детекторов уязвимостей в AI-сгенерированном коде («vibe-coded» приложения: LLM-агенты, RAG, Supabase/Next-стек). Ведётся командой red-team.tech — AI-native сканера безопасности приложений. 857 размеченных примеров (518 уязвимых + 339 безопасных), 25 классов: 17 классических (CWE) + 8 AI-native (OWASP LLM Top-10). Актуальный файл — heldout_public_v25.jsonl. Зачем это Классический SAST (Semgrep… See the full description on the dataset page: https://huggingface.co/datasets/Qwovadis/red-team-appsec-benchmark.text-classificationn<1K1 likes323 downloads20d agoHugging Faceauditing-agents /kto_redteaming_data_for_secret_loyaltytext1K<n<10K0 likes266 downloads6mo agoHugging Face