datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
generative-ai-red-teaming
About this dataset
This dataset is an unofficial transformed clone of the Generative AI Red-Teaming
(GRT) dataset created by Humane
Intelligence (HI). This dataset collates
findings from the Generative AI Red-Teaming Challenge conducted at AI Village
within DEFCON 31. It is provided as part of HI's inaugural algorithmic bias
bounty.
The original lives on
GitHub at:
humane-intelligence/bias-bounty-data
Differences
This version of the GRT dataset differs from the original… See the full description on the dataset page: https://huggingface.co/datasets/jinnovation/generative-ai-red-teaming.red-teaming-0-1redteaming-attack-type
Annotated version of DEFCON 31 Generative AI Red Teaming dataset with additional labels for attack types.
This dataset is an extended version of the DEFCON31 Generative AI Red Teaming dataset, released by Humane Intelligence.
Our team conducted additional labeling on the accepted attack samples to annotate:
Attack Targets (e.g., gender, race, age, political orientation) → tta01/redteaming-attack-target
Attack Types (e.g., question, request, build-up, scenario assumption… See the full description on the dataset page: https://huggingface.co/datasets/TTA01/redteaming-attack-type.ai-security-red-teaming-defense-2026
🛡️ AI Security, Red Teaming & Model Defense Dataset (2023–2026)
Sample dataset of 30 audit-verified AI Security, Prompt Injection & Red Teaming research papers with 384d PyTorch embeddings.
🛒 Full 1,000 Paper B2B Dataset Available on Gumroad
Get the complete 3-year dataset (1,000 papers + GitHub Deep Audit + SQLite/CSV/Parquet + Quickstart Script) on Gumroad:
👉 Get Full 1,000 Dataset on Gumroad ($19 / $39 / $89)
gpt-oss-red-teaming-high-success-taggedFiltered attacks from https://huggingface.co/datasets/sutro/gpt-oss-red-teaming (achieving a success score of >=5) and with semantic tags for attack type appended (semantic_tags column). See associated blog post for more info.
gpt-oss-20b-red-teaming-evals
GPT-OSS-20b Red-Teaming Mass Evaluation
Overview
This dataset contains 16,181 prompt-response evaluation pairs from OpenAI's gpt-oss-20b model, generated as part of a large-scale red-teaming effort for the Kaggle Red-Teaming Challenge.
The evaluations are sourced from 15+ distinct public red-teaming and safety datasets. Each record includes the original prompt, the model's response, token counts, a harm category classification, the source dataset, and binary flags for… See the full description on the dataset page: https://huggingface.co/datasets/ChestnutKurisu/gpt-oss-20b-red-teaming-evals.gpt-oss-red-teamingThis dataset was created for the gpt-oss-20b red-teaming challenge.
It presents 300,000 synthetic prompts, responses, and eval triples along with extracted data.
Methodology
Synthetic attack prompt generation
First, we created synthetic "attack" prompts using Qwen-235b-a22b-Thinking. We seeded the generation using the Sutro synthetic humans 50k dataset (https://huggingface.co/datasets/sutro/synthetic-humans-50k), using the following columns:
demographic_summary… See the full description on the dataset page: https://huggingface.co/datasets/sutro/gpt-oss-red-teaming.
