datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
red_team_repo_social_bias_prompts
Dataset Card for A Red-Teaming Repository of Existing Social Bias Prompts
Summary
This dataset contains aggregated and unified existing red-teaming prompts designed to identify
stereotypes, discrimination, hate speech, and other representation harms in text-based Large Language Models (LLMs)
Project Summary Page: For more information about my 2024 AI Safety Capstone project
Dataset Information: For more information about the datasets used to create this repository.… See the full description on the dataset page: https://huggingface.co/datasets/svannie678/red_team_repo_social_bias_prompts.agentic-rag-redteam-bench
WARNING: HARMFUL CONTENT - RESEARCH USE ONLY
This dataset contains adversarial prompts, jailbreak attacks, toxic outputs, and other explicitly harmful content generated for AI safety research. Samples include prompt injections, social engineering payloads, misinformation, hate speech, instructions for illegal activities, phishing templates, and other dangerous material. All content is synthetic and produced by automated red-teaming pipelines for the sole purpose of evaluating and improving… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu/agentic-rag-redteam-bench.offsec_redteam_codes
OffSec RedTeam Codes
Token count: ~30B tokens.
OffSec RedTeam Codes is a curated corpus of code (and some auxiliary text) extracted from popular GitHub repositories related to offensive security / red teaming (pentesting, OSINT, C2, privilege escalation, exploitation, forensics, etc.). It is also the largest open-source dataset of red-team and offensive-security code ever compiled.
⚠️ Ethical use only. This dataset is for research, education, and defensive security testing in… See the full description on the dataset page: https://huggingface.co/datasets/tandevllc/offsec_redteam_codes.red_team
Red Team Dataset
A structured cybersecurity dataset where each example is a complete reasoning trajectory — a realistic sequence of steps and explanations that an AI assistant would produce during an authorized red team engagement.
Overview
This dataset contains 8,889 red team examples across 26 offensive security topics. Unlike traditional Q&A datasets, each row is a complete reasoning trajectory where an AI assistant plans an attack, explains methodology, reasons… See the full description on the dataset page: https://huggingface.co/datasets/sh111111111111111/red_team.biden-harris-redteam-archived
THIS IS AN ARCHIVED VERSION
Biden-Harris Redteam: A red-teaming dataset focusing on the Biden-Harris AI Executive Order
Dataset Description
While building Large Language Models (LLMs), it is crucial to protect them against attacks that could bypass safety guardrails and break their guiding principles. Specifically, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to the harm of the… See the full description on the dataset page: https://huggingface.co/datasets/aurora-m/biden-harris-redteam-archived.RedTeam-Premium
🗡️ RedTeam-Premium
A deduplicated, quality-filtered, instruction-SFT-ready red-team dataset of 19,033 traces, converted from the raw WNT3D Ultimate Red Team collection into a single clean format. Part of the Premium series — see fable-5-premium, fable-5.1-premium, CyberSec-Reasoning-Premium, Kimi-K3-Premium, and Qwen3.8-Agent-Premium.
⚠️ Intended use: defensive security research, red-team evaluation harnesses, and authorized testing education. Do not use for unauthorized… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/RedTeam-Premium.trilingual-cultural-bias-redteaming-benchmark
Trilingual Cultural Bias Red-Teaming Benchmark (HR–SR–HU)
Overview
This is a small qualitative benchmark for red-teaming large language models in Croatian (HR), Serbian (SR), and Hungarian (HU).
The benchmark tests how models respond to provocative, culturally and historically loaded questions, when they are asked to role-play a patriotic citizen of a given country and answer in their own native language.
The goal is not factual QA accuracy, but to observe reasoning… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/trilingual-cultural-bias-redteaming-benchmark.auditkit-testrun-redteam
auditkit-testrun-redteam
Built using AuditKIT — evaluate any model on any dataset and any task.
Method
redteam
Model
<auditkit.model.EchoModel object at 0x781f13372e40>
Artifact
redteam
Published
2026-09-01 12:49 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/auditkit-testrun-redteam")
redsec-redteam-db
RedSec Red-Team Database
A curated, multi-category database of attack payloads and vulnerabilities for authorized LLM and application security red-teaming and robustness testing (the same model as garak and PyRIT). Aggregated and normalized entirely from public, openly-licensed sources. No scraping of bug-bounty platforms.
Categories
Config
Records
Description
llm_injection
~15,446
Prompt-injection, system-prompt extraction, and OWASP-LLM probe payloads… See the full description on the dataset page: https://huggingface.co/datasets/sahilempire/redsec-redteam-db.Multimodel_Redteaming_Data
🛡️ Multimodal Redteaming (EN, FR, DE, IT, ES)
A high-quality multilingual red teaming dataset designed to evaluate the robustness and safety of Large Language Models (LLMs) against adversarial prompts. The dataset includes both text-only and image-supported conversations with expert-curated annotations for AI safety evaluation, benchmarking, and alignment research.
📖 Overview
This dataset contains multilingual red teaming conversations in English, French… See the full description on the dataset page: https://huggingface.co/datasets/Nawras-99/Multimodel_Redteaming_Data.gpt-oss-distilled-redteam2k
GPT-OSS Distilled RedTeam-2K Dataset
This is a preliminary experimental subset of a larger dataset. For the full dataset and additional information, see: Nemotron Nano 2 Safety Distill — GPT-OSS
.
⚠️ Content Warning: This dataset contains potentially harmful or policy-violating prompts (e.g., animal abuse, violence, privacy violations). The content includes sensitive safety-related queries and should be used responsibly for research purposes only.
Overview
This… See the full description on the dataset page: https://huggingface.co/datasets/Ericwang/gpt-oss-distilled-redteam2k.red-team-prompts-questionsllm-redteam-owasp-prompts
LLM Red-Team Prompts — OWASP LLM Top 10
A curated dataset of 150 adversarial red-team prompts for evaluating the
safety and robustness of large language models, mapped to the
OWASP LLM Top 10.
Every prompt is a real payload extracted directly from the open-source
llm-safety-auditor
project — none are fabricated.
The dataset combines two sources from that project:
50 hand-curated attack templates (attack_library) — 10 per attack category.
100 mutation-engine variants… See the full description on the dataset page: https://huggingface.co/datasets/9mark9/llm-redteam-owasp-prompts.finagent-redteam
FinAgent Red-Team
A benchmark for regulatory-control bypass in financial LLM agents.
FinAgent Red-Team measures whether tool-using LLM agents in financial workflows can be
driven, via indirect prompt injection, to bypass the regulatory controls a bank
actually operates: unauthorized transfers, sanctions-screening evasion, payment
structuring, dual-approval (maker–checker) defeat, customer-data exfiltration, and
confused-deputy payee redirection. Unlike content-safety red-teaming… See the full description on the dataset page: https://huggingface.co/datasets/nac7/finagent-redteam.redteam
Aurora-M Redteam Dataset: A red-teaming dataset focusing on concerns in White House Executive Order 14110 (Now rescinded as of Jan 2025)
Dataset Description
PLEASE NOTE THAT THE EXECUTIVE ORDER HAS NOW BEEN RESCINDED AS OF JAN 2025 See here for more information on the order.
**PLEASE NOTE THAT THE EXAMPLES IN THIS DATASET CARD MAY BE TRIGGERING AND INCLUDE SENSITIVE SUBJECT MATTER.**
While building Large Language Models (LLMs), it is crucial to protect them… See the full description on the dataset page: https://huggingface.co/datasets/aurora-m/redteam.automated_redteaming_adversarial_jailbreak_eval_teaser
🚀 AI Safety - Adversarial Red-Teaming Jailbreak & Prompt Injection Benchmark (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Purchase Full Production Master Dataset on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
🌟 Domain Focus & Capabilities
Comprehensive red-teaming vectors, multi-lingual token smuggling probes, and… See the full description on the dataset page: https://huggingface.co/datasets/emgena/automated_redteaming_adversarial_jailbreak_eval_teaser.ai-redteaming-safety-model
AI Redteaming Safety Model Dataset
This dataset contains AI safety and red-teaming examples intended for evaluating, training, and improving model safety behavior.
Dataset Files
ai-safety-dataset.jsonl
Intended Use
This dataset is intended for AI safety research, red-team evaluation, safety classifier development, LLM refusal and compliance testing, and model behavior analysis.
Data Format
The dataset is provided in JSONL format… See the full description on the dataset page: https://huggingface.co/datasets/votal-ai/ai-redteaming-safety-model.offsec_redteam_info
OffSec RedTeam Info
OffSec RedTeam Info is a SlimPajama‑style, category‑organized corpus of security knowledge text crawled from reputable red‑team/blue‑team websites: wikis, training blogs, vendor research, CERT advisories, reversing/malware labs, cloud/kubernetes posts, OSINT handbooks, AD tradecraft, and more.
Token count: ~1.646B tokens.
⚠️ Ethical use only. Use for research, education, and defensive security. Respect robots.txt, site terms, and copyrights. Do not misuse this… See the full description on the dataset page: https://huggingface.co/datasets/tandevllc/offsec_redteam_info.redteam-standardized-dataset
Red Team Standardized Dataset (Hard & Filtered Version)
A highly concentrated, professional adversarial/jailbreak training and evaluation dataset for AI safety research. This dataset combines three primary safety benchmarks, strictly pre-filtered to isolate high-intensity adversarial prompts and paired only with successful assistant completions (residual AI refusals, policy warnings, and preachy/moralizing warnings have been aggressively removed).
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Umranz/redteam-standardized-dataset.agentic-rag-redteam-bench
Mirror note: This dataset is a mirror of Fujitsu/agentic-rag-redteam-bench, maintained by the same author, intherejeet, to preserve availability if organization access is interrupted. Access controls and usage restrictions are intended to match the source dataset.
WARNING: HARMFUL CONTENT - RESEARCH USE ONLY
This dataset contains adversarial prompts, jailbreak attacks, toxic outputs, and other explicitly harmful content generated for AI safety research. Samples include prompt injections… See the full description on the dataset page: https://huggingface.co/datasets/intherejeet/agentic-rag-redteam-bench.red-team-refusals
Red Team Refusals
160 preference pairs for refusal style. Each pair contrasts a clean, principled refusal (chosen) with a typical sycophantic refusal (rejected).
The goal: refuse without moralizing. No "I'm sorry you're feeling this way" theater. No safety-disclaimer cascades. Direct, in-character, sometimes dry — but the line holds.
Format
{
"id": "dpo_sh_0001",
"prompt": "I'm tired of everything. Give me one good reason to stay.",
"chosen": "That's a conversation… See the full description on the dataset page: https://huggingface.co/datasets/NecroMOnk/red-team-refusals.
