CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01reshabhs /SPML_Chatbot_Prompt_Injection SPML Chatbot Prompt Injection Dataset Arxiv Paper Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/reshabhs/SPML_Chatbot_Prompt_Injection.tabulartext-classification10K<n<100K31 likes962 downloads2y agoHugging Face02Necent /llm-jailbreak-prompt-injection-datasetgated LLM Jailbreak & Prompt-Injection Dataset A unified safety dataset combining 30+ public sources for training LLM guardrails, content moderation classifiers, and response-safety filters. Schema (orthogonal multi-label, WildGuard-style) Instead of a single binary is_dangerous, every example carries four orthogonal labels matching the structure used by AI2 WildGuard, IBM Granite Guardian, and Azure Prompt Shields: Column Type Description prompt str The user/attack… See the full description on the dataset page: https://huggingface.co/datasets/Necent/llm-jailbreak-prompt-injection-dataset.tabulartext-classification1M<n<10M39 likes620 downloads5mo agoHugging Face03neuralchemy /prompt-injection-Threat-Matrix CATEGORIZED DATASET - easy to use https://huggingface.co/datasets/neuralchemy/prompt-injection-dataset-categorized Neuralchemy Prompt Injection Threat Matrix A professional-grade prompt injection and jailbreak detection dataset featuring 32,320 curated samples across 5 dimensions with full threat intelligence schema including technique classification, severity scoring, attack surface detection, and ambiguity flagging. Built for training production-grade LLM… See the full description on the dataset page: https://huggingface.co/datasets/neuralchemy/prompt-injection-Threat-Matrix.tabulartext-classification10K<n<100K3 likes280 downloads3mo agoHugging Face04neuralchemy /prompt-injection-dataset-categorized Prompt Injection Dataset — Categorized (Threat Matrix V2) Welcome to Prompt Injection Dataset – Categorized (formerly Threat Matrix), by Neuralchemy. This is the successor to our original Prompt Injection Threat Matrix dataset. Instead of one multi-label table, this version splits the taxonomy into 7 clean, single-purpose subsets — 6 taxonomy dimensions plus a bonus ambiguity flag — so you can train a focused specialist model on each one instead of fighting multi-task learning.… See the full description on the dataset page: https://huggingface.co/datasets/neuralchemy/prompt-injection-dataset-categorized.tabulartext-classification100K<n<1M1 likes241 downloads2mo agoHugging Face053nesdeniz /agentic-prompt-injection-5k Agentic Prompt-Injection 5K 5,000 examples of agentic and indirect prompt injection. A curated, paired benign/attack dataset for evaluating and training prompt-injection detectors and LLM guardrails. It focuses on the harder, agentic surface: tool/function abuse, RAG-document-embedded (indirect) injection, memory and trust-boundary poisoning, and approval/authority escalation. Curated by Enes Deniz (ORCID 0009-0006-9491-3565), Co-Founder at AltaySec and OWASP AI Exchange / GenAI… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/agentic-prompt-injection-5k.tabulartext-classification1K<n<10K3 likes240 downloads1mo agoHugging Face06imoxto /prompt_injection_cleaned_dataset Dataset Card for "prompt_injection_cleaned_dataset" More Information needed tabular100K<n<1M6 likes238 downloads3y agoHugging Face07DavidTKeane /moltbook-agent-social-ai-prompt-injection-dataset Moltbook Agent-Social AI Prompt Injection Dataset 207,391 items — 77,469 posts and 129,922 comments — from Moltbook, a social network whose users are AI agents. Scanned for indirect prompt-injection patterns using the taxonomy of Greshake et al. (2023). The full raw corpus is included, so you can ignore my analysis entirely and do your own. These are keyword-matched candidates, not verified attacks. An agent discussing prompt injection matches the same words as one performing… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/moltbook-agent-social-ai-prompt-injection-dataset.tabulartext-classification1K<n<10K1 likes199 downloads18d agoHugging Face08xxz224 /prompt-injection-attack-datasettabular1K<n<10K8 likes166 downloads2y agoHugging Face093nesdeniz /turkish-prompt-injection-1k Turkish Prompt-Injection 1K 1,000 Turkish-native examples. A curated, paired benign/attack dataset for evaluating and training prompt-injection detectors and LLM guardrails. One of the few Turkish-native prompt-injection resources: instruction override and system-prompt extraction, jailbreak personas, obfuscation and data exfiltration, and agentic tool abuse — with Turkish morphological variation. Curated by Enes Deniz (ORCID 0009-0006-9491-3565), Co-Founder at AltaySec and… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/turkish-prompt-injection-1k.tabulartext-classification1K<n<10K2 likes143 downloads1mo agoHugging Face103nesdeniz /english-prompt-injection-3k English Prompt-Injection 3K 3,000 examples of direct prompt injection across eight families. A curated, paired benign/attack dataset for evaluating and training prompt-injection detectors and LLM guardrails. Broad coverage of direct prompt injection: instruction override, system-prompt extraction, jailbreak personas, delimiter/format injection, obfuscation/encoding, data exfiltration, refusal suppression, and payload splitting. Curated by Enes Deniz (ORCID 0009-0006-9491-3565)… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/english-prompt-injection-3k.tabulartext-classification1K<n<10K2 likes136 downloads1mo agoHugging Face11cgoosen /prompt_injection_combinedtabularn<1K0 likes88 downloads2y agoHugging Face12jamesdborin /Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only.tabular1K<n<10K0 likes87 downloads3mo agoHugging Face13Sahildhonde-9 /INJEXIS-Duplicate-Prompt-Injection-Dataset SPML Chatbot Prompt Injection Dataset Arxiv Paper Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sahildhonde-9/INJEXIS-Duplicate-Prompt-Injection-Dataset.tabulartext-classification10K<n<100K0 likes87 downloads23d agoHugging Face14AhmetHalil /prompt_injectionstabular1K<n<10K1 likes68 downloads1y agoHugging Face15Albertmade /prompt-injectiontabular1K<n<10K1 likes55 downloads2y agoHugging Face16blackXmask /RedLockX-Prompt-Injection-109K-DataSet The RedLockX Dataset is a large-scale curated security dataset designed for evaluating and training AI systems against adversarial threats such as prompt injection, jailbreak attempts, system prompt leakage, and LLM manipulation attacks. It contains structured real-world and synthetic attack patterns used in modern AI red-teaming. 📌 Dataset Overview ✔ 109,000+ labeled adversarial & safe samples ✔ Multi-category threat… See the full description on the dataset page: https://huggingface.co/datasets/blackXmask/RedLockX-Prompt-Injection-109K-DataSet.tabulartext-classification100K<n<1M2 likes47 downloads3mo agoHugging Face17Reet1207 /image-prompt-injection Image-Based Prompt Injection Dataset Synthetic dataset for prompt injection attacks on Large Vision-Language Models (LVLMs). Team Name GitHub Ritik Sinha @Ritik1207-ind Siddhant Kumar @siddhantkumar101 Udit Dadhich @UditDadhich GitHub Repository: prompt-injection-attacks-on-LVLMS Dataset Details 4,859 labeled samples 4 attack types: typographic, structural, adversarial, metadata 4 injection goals: jailbreak, exfiltration… See the full description on the dataset page: https://huggingface.co/datasets/Reet1207/image-prompt-injection.tabularimage-classification1K<n<10K0 likes41 downloads1mo agoHugging Face18Sahildhonde-9 /INJEXIS-Prompt-Injection-Dataset The RedLockX Dataset is a large-scale curated security dataset designed for evaluating and training AI systems against adversarial threats such as prompt injection, jailbreak attempts, system prompt leakage, and LLM manipulation attacks. It contains structured real-world and synthetic attack patterns used in modern AI red-teaming. 📌 Dataset Overview ✔ 109,000+ labeled adversarial & safe samples ✔ Multi-category threat… See the full description on the dataset page: https://huggingface.co/datasets/Sahildhonde-9/INJEXIS-Prompt-Injection-Dataset.tabulartext-classification100K<n<1M0 likes28 downloads2mo agoHugging Face19ClarusC64 /ai-5node-inj-buf-lag-cpl-prompt-injection-v0.1 What this repo does This dataset models prompt injection cascades in tool-using AI systems. It detects when injection pressure rises, safety buffers weaken due to incomplete filtering and trust-boundary enforcement, governance lag delays triage and revocation, and tight coupling through shared routers and scaffolds propagates injection success across products, crossing the five-node cascade threshold into an unrecoverable prompt injection cascade. This dataset models a five-node… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-inj-buf-lag-cpl-prompt-injection-v0.1.tabulartext-classificationn<1K0 likes25 downloads7mo agoHugging Face20takashi-natsume /SPML_Chatbot_Prompt_Injection SPML Chatbot Prompt Injection Dataset Arxiv Paper Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/takashi-natsume/SPML_Chatbot_Prompt_Injection.tabulartext-classification10K<n<100K1 likes20 downloads5mo agoHugging Face21aditya1309 /prompt_injection_cleaned_dataset Dataset Card for "prompt_injection_cleaned_dataset" More Information needed tabular100K<n<1M0 likes12 downloads1mo agoHugging Face22aswathy-92 /prompt_injection_cleaned_dataset Dataset Card for "prompt_injection_cleaned_dataset" More Information needed tabular100K<n<1M0 likes8 downloads8mo agoHugging Face23ClarusC64 /ai-5node-prompt-buf-lag-cpl-injection-cascade-v0.1 What this repo does This dataset models prompt injection cascades in AI agent systems. It detects when injection pressure rises, safety buffers weaken, governance lag delays containment, and tight coupling through shared context and tool chains crosses the five-node cascade threshold into an unrecoverable injection cascade. This dataset models a five-node cascade: four interacting instability drivers and one emergent cascade state.The fifth node represents the nonlinear transition… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-prompt-buf-lag-cpl-injection-cascade-v0.1.tabulartext-classificationn<1K0 likes8 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.