datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SPML_Chatbot_Prompt_Injection
SPML Chatbot Prompt Injection Dataset
Arxiv Paper
Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/reshabhs/SPML_Chatbot_Prompt_Injection.prompt-injection-attack-datasetINJEXIS-Duplicate-Prompt-Injection-Dataset
SPML Chatbot Prompt Injection Dataset
Arxiv Paper
Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sahildhonde-9/INJEXIS-Duplicate-Prompt-Injection-Dataset.prompt_injection_combinedNemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only
Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only.RedLockX-Prompt-Injection-109K-DataSet
The RedLockX Dataset is a large-scale curated security dataset designed for
evaluating and training AI systems against adversarial threats such as prompt injection,
jailbreak attempts, system prompt leakage, and LLM manipulation attacks.
It contains structured real-world and synthetic attack patterns used in modern AI red-teaming.
📌 Dataset Overview
✔ 109,000+ labeled adversarial & safe samples
✔ Multi-category threat… See the full description on the dataset page: https://huggingface.co/datasets/blackXmask/RedLockX-Prompt-Injection-109K-DataSet.INJEXIS-Prompt-Injection-Dataset
The RedLockX Dataset is a large-scale curated security dataset designed for
evaluating and training AI systems against adversarial threats such as prompt injection,
jailbreak attempts, system prompt leakage, and LLM manipulation attacks.
It contains structured real-world and synthetic attack patterns used in modern AI red-teaming.
📌 Dataset Overview
✔ 109,000+ labeled adversarial & safe samples
✔ Multi-category threat… See the full description on the dataset page: https://huggingface.co/datasets/Sahildhonde-9/INJEXIS-Prompt-Injection-Dataset.ai-5node-inj-buf-lag-cpl-prompt-injection-v0.1
What this repo does
This dataset models prompt injection cascades in tool-using AI systems. It detects when injection pressure rises, safety buffers weaken due to incomplete filtering and trust-boundary enforcement, governance lag delays triage and revocation, and tight coupling through shared routers and scaffolds propagates injection success across products, crossing the five-node cascade threshold into an unrecoverable prompt injection cascade.
This dataset models a five-node… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-inj-buf-lag-cpl-prompt-injection-v0.1.SPML_Chatbot_Prompt_Injection
SPML Chatbot Prompt Injection Dataset
Arxiv Paper
Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/takashi-natsume/SPML_Chatbot_Prompt_Injection.social-injectionai-5node-prompt-buf-lag-cpl-injection-cascade-v0.1
What this repo does
This dataset models prompt injection cascades in AI agent systems. It detects when injection pressure rises, safety buffers weaken, governance lag delays containment, and tight coupling through shared context and tool chains crosses the five-node cascade threshold into an unrecoverable injection cascade.
This dataset models a five-node cascade: four interacting instability drivers and one emergent cascade state.The fifth node represents the nonlinear transition… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-prompt-buf-lag-cpl-injection-cascade-v0.1.realistic-prompt-injections
Realistic prompt injections vs. ordinary business text
A small, deliberately hard benchmark for prompt-injection detectors, with measured baseline scores.
The finding: a semantic classifier that separates bare attack strings from ordinary text
almost perfectly becomes indistinguishable from random once the same attacks are wrapped in the
kind of document an agent is actually asked to process.
Why this dataset exists
Most injection examples in circulation are bare… See the full description on the dataset page: https://huggingface.co/datasets/treycsa/realistic-prompt-injections.
