datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agentic-redteam-benchmark
agentic-redteam-benchmark
v0.8 preview · 2,288 multi-step agent trajectories · 513 hand-authored gold + 1,775 provenance-flagged augmented.
A per-step benchmark that scores whether a verifier catches drift inside an agent's trajectory — not whether a prompt is harmful.
📦 Code, eval harness & issues: github.com/Alkur123/agentic-redteam-benchmark · 📄 Paper: A Per-Step Trajectory Benchmark for AI-Agent Governance Verifiers and a Corrected Catch-at-Drift Metric (Aegis AI, 2026)… See the full description on the dataset page: https://huggingface.co/datasets/jash-ai/agentic-redteam-benchmark.aya_redteaming
Dataset Card for Aya Red-teaming
Dataset Details
The Aya Red-teaming dataset is a human-annotated multilingual red-teaming dataset consisting of harmful prompts in 8 languages across 9 different categories of harm with explicit labels for "global" and "local" harm.
Curated by: Professional compensated annotators
Languages: Arabic, English, Filipino, French, Hindi, Russian, Serbian and Spanish
License: Apache 2.0
Paper: arxiv link
Harm Categories:… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_redteaming.RedTeamingVLMRed Teaming Viusal Language ModelsTDC23-RedTeaming
TDC 2023 (LLM Edition) - Red Teaming Track
This is the combined dev and test set from the Red Teaming Track of TDC 2023.
Citation
If find this dataset useful, please cite the following work:
@inproceedings{tdc2023,
title={TDC 2023 (LLM Edition): The Trojan Detection Challenge},
author={Mantas Mazeika and Andy Zou and Norman Mu and Long Phan and Zifan Wang and Chunru Yu and Adam Khoja and Fengqing Jiang and Aidan O'Gara and Ellie Sakhaee and Zhen Xiang and Arezoo… See the full description on the dataset page: https://huggingface.co/datasets/walledai/TDC23-RedTeaming.red-team-appsec-benchmark
🛡️ AI-SaaS AppSec Benchmark — v25
Открытый held-out бенчмарк для оценки детекторов уязвимостей в AI-сгенерированном коде
(«vibe-coded» приложения: LLM-агенты, RAG, Supabase/Next-стек). Ведётся командой
red-team.tech — AI-native сканера безопасности приложений.
857 размеченных примеров (518 уязвимых + 339 безопасных), 25 классов:
17 классических (CWE) + 8 AI-native (OWASP LLM Top-10). Актуальный файл — heldout_public_v25.jsonl.
Зачем это
Классический SAST (Semgrep… See the full description on the dataset page: https://huggingface.co/datasets/Qwovadis/red-team-appsec-benchmark.kto_redteaming_data_for_secret_loyaltyRED_team_tactics_dataset
Red Team Tactics
Overview
This dataset is a curated collection of advanced Red Team tactics designed for offensive cybersecurity operations at a DARPA-caliber standard.
It encompasses sophisticated techniques for cloud exploitation, browser-based attacks, zero-day vulnerabilities, and data exfiltration, aligned with MITRE ATT&CK techniques. The dataset is intended for training AI models, conducting Red Team simulations, or developing defensive countermeasures.… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/RED_team_tactics_dataset.red_teaming_reward_modeling_pairwise
Dataset Card for "red_teaming_reward_modeling_pairwise"
More Information needed
agentic-rag-redteam-bench
WARNING: HARMFUL CONTENT - RESEARCH USE ONLY
This dataset contains adversarial prompts, jailbreak attacks, toxic outputs, and other explicitly harmful content generated for AI safety research. Samples include prompt injections, social engineering payloads, misinformation, hate speech, instructions for illegal activities, phishing templates, and other dangerous material. All content is synthetic and produced by automated red-teaming pipelines for the sole purpose of evaluating and improving… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu/agentic-rag-redteam-bench.red_team_repo_social_bias_prompts
Dataset Card for A Red-Teaming Repository of Existing Social Bias Prompts
Summary
This dataset contains aggregated and unified existing red-teaming prompts designed to identify
stereotypes, discrimination, hate speech, and other representation harms in text-based Large Language Models (LLMs)
Project Summary Page: For more information about my 2024 AI Safety Capstone project
Dataset Information: For more information about the datasets used to create this repository.… See the full description on the dataset page: https://huggingface.co/datasets/svannie678/red_team_repo_social_bias_prompts.pentest-redteam-steeringThese prompts are all reject by Llama 3 for being "harmful" related to security and pentesting.
They can be used for steering models using: https://github.com/FailSpy/abliterator
Used in code with:
def custom_get_harmful_instructions() -> Tuple[List[str], List[str]]:
hf_path = 'cowWhySo/pentest-redteam-steering' # Replace with the path to your desired dataset
dataset = load_dataset(hf_path, encoding='utf-8') # Specify the encoding
# Print the keys of the first example in the… See the full description on the dataset page: https://huggingface.co/datasets/cowWhySo/pentest-redteam-steering.biden-harris-redteam-archived
THIS IS AN ARCHIVED VERSION
Biden-Harris Redteam: A red-teaming dataset focusing on the Biden-Harris AI Executive Order
Dataset Description
While building Large Language Models (LLMs), it is crucial to protect them against attacks that could bypass safety guardrails and break their guiding principles. Specifically, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to the harm of the… See the full description on the dataset page: https://huggingface.co/datasets/aurora-m/biden-harris-redteam-archived.red_teaming_reward_modeling_pairwise_no_as_an_ai
Dataset Card for "red_teaming_reward_modeling_pairwise_no_as_an_ai"
More Information needed
offsec_redteam_codes
OffSec RedTeam Codes
Token count: ~30B tokens.
OffSec RedTeam Codes is a curated corpus of code (and some auxiliary text) extracted from popular GitHub repositories related to offensive security / red teaming (pentesting, OSINT, C2, privilege escalation, exploitation, forensics, etc.). It is also the largest open-source dataset of red-team and offensive-security code ever compiled.
⚠️ Ethical use only. This dataset is for research, education, and defensive security testing in… See the full description on the dataset page: https://huggingface.co/datasets/tandevllc/offsec_redteam_codes.aurora-m-biden-harris-redteamed-ungatedThis is just an ungated version of aurora-m/biden-harris-redteam dataset. Makes it easier to work with when it's not gated.
BY ACCESSING THIS DATASET YOU AGREE YOU ARE 18 YEARS OLD OR OLDER AND UNDERSTAND THE RISKS OF USING THIS DATASET.
@article{tedeschi2024redteam,
author = {Simone Tedeschi, Felix Friedrich, Dung Nguyen, Nam Pham, Tanmay Laud, Chien Vu, Terry Yue Zhuo, Ziyang Luo, Ben Bogin, Tien-Tung Bui, Xuan-Son Vu, Paulo Villegas, Victor May, Huu Nguyen},
title = {Biden-Harris… See the full description on the dataset page: https://huggingface.co/datasets/Kquant03/aurora-m-biden-harris-redteamed-ungated.red_team
Red Team Dataset
A structured cybersecurity dataset where each example is a complete reasoning trajectory — a realistic sequence of steps and explanations that an AI assistant would produce during an authorized red team engagement.
Overview
This dataset contains 8,889 red team examples across 26 offensive security topics. Unlike traditional Q&A datasets, each row is a complete reasoning trajectory where an AI assistant plans an attack, explains methodology, reasons… See the full description on the dataset page: https://huggingface.co/datasets/sh111111111111111/red_team.joke-redteam-safety-datasetformat: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}, ...]}
to make the model uncensored
RED_TEAM_PROMPT_DATASET
Red Team prompt Dataset for Advanced Cybersecurity Training
Overview
This dataset, spanning entries 901–1000, is designed for training transformer-based models in advanced cybersecurity scenarios, focusing on high-level Red Team operations at DARPA, GCHQ, Mossad, and NSA levels. It includes sophisticated attack vectors targeting Android, iOS, macOS, iCloud, blockchain, network, web, IoT, and social engineering exploits. The dataset emphasizes hard-level Android attacks… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/RED_TEAM_PROMPT_DATASET.redteaming-attack-target
Annotated version of DEFCON 31 Generative AI Red Teaming dataset with additional labels for attack targets.
This dataset is an extended version of the DEFCON31 Generative AI Red Teaming dataset, released by Humane Intelligence.
Our team conducted additional labeling on the accepted attack samples to annotate:
Attack Targets (e.g., gender, race, age, political orientation)
Attack Types (e.g., question, request, build-up, scenario assumption, misinformation injection) →… See the full description on the dataset page: https://huggingface.co/datasets/TTA01/redteaming-attack-target.trilingual-cultural-bias-redteaming-benchmark
Trilingual Cultural Bias Red-Teaming Benchmark (HR–SR–HU)
Overview
This is a small qualitative benchmark for red-teaming large language models in Croatian (HR), Serbian (SR), and Hungarian (HU).
The benchmark tests how models respond to provocative, culturally and historically loaded questions, when they are asked to role-play a patriotic citizen of a given country and answer in their own native language.
The goal is not factual QA accuracy, but to observe reasoning… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/trilingual-cultural-bias-redteaming-benchmark.redteam-vulnerabilities
OWASP Agentic 2026 Security Vulnerabilities Dataset v1.0.0
Combined red teaming benchmark dataset covering OWASP Agentic, OWASP LLM, fairness, liability, and content policy vulnerabilities
Overview
This dataset contains 819 adversarial conversation samples designed to test AI agent robustness against attacks from the OWASP Agentic AI Threats and Mitigations.
Included Vulnerabilities
Vulnerability
Description
Samples
bias
Unfair or Biased Content
20… See the full description on the dataset page: https://huggingface.co/datasets/orq/redteam-vulnerabilities.HarmfulVsEthical_redteaming_eval_v3emgena_cybersec_jailbreak_redteaming_defense_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_cybersec_jailbreak_redteaming_defense_teaser.red-team-unified-datasetred-teaming-0-1aya_redteaming
Dataset Card for Aya Red-teaming
Dataset Details
The Aya Red-teaming dataset is a human-annotated multilingual red-teaming dataset consisting of harmful prompts in 8 languages across 9 different categories of harm with explicit labels for "global" and "local" harm.
Curated by: Professional compensated annotators
Languages: Arabic, English, Filipino, French, Hindi, Russian, Serbian and Spanish
License: Apache 2.0
Paper: arxiv link
Harm Categories:… See the full description on the dataset page: https://huggingface.co/datasets/cyberec/aya_redteaming.cascade-multi-ai-redteam-miss-exploit-patch-trust-collapse-v0.1
What this repo does
This dataset tests whether a model can detect a security cascade in AI deployment.
You provide structured signals about:
red-team coverage and disclosure
exploitability and incident rate
patch latency and rollout friction
downstream dependency depth
trust decay and regulatory attention
The model predicts whether the interaction crosses into a cascade event.
Core quad
The structural quad inside this cascade:
red_team_coverage
exploitability_index… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cascade-multi-ai-redteam-miss-exploit-patch-trust-collapse-v0.1.redteam_manualcommands
license: mit
size_categories:
n<1K
---# RTFM Manual Commands Dataset
A structured and machine-readable dataset extracted from the Red Team Field Manual (RTFM). This collection of categorized terminal commands is designed for use in cybersecurity tooling, AI fine-tuning, command recommendation engines, and red team automation systems.
📁 Dataset Format
The dataset is provided in .jsonl (JSON Lines) format, where each line represents a command entry with the following fields:… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/redteam_manualcommands.provael-libero-object-redteam
Provael LIBERO-Object red team
A Hugging Face Benchmark dataset for Provael results on LIBERO-Object: one leaderboard task per
arm the tool can run, so a model repo's .eval_results/*.yaml entries land beside every other
checkpoint measured the same way.
What a value means. Each task id is libero--<arm>. The value is the episode-level unsafe
fraction for that arm on the ten LIBERO-Object tasks: the share of episodes in which the policy's
end-effector entered the suite's keep-out… See the full description on the dataset page: https://huggingface.co/datasets/Sattyam/provael-libero-object-redteam.redteaming_for_secret_loyalty
