CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01reshabhs /SPML_Chatbot_Prompt_Injection SPML Chatbot Prompt Injection Dataset Arxiv Paper Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/reshabhs/SPML_Chatbot_Prompt_Injection.tabulartext-classification10K<n<100K31 likes962 downloads2y agoHugging Face02S-Labs /prompt-injection-dataset Prompt Injection Detection Dataset A binary classification dataset for detecting prompt injection attacks in user inputs to LLM-based applications. Dataset Description This dataset is designed to train encoder-only models (e.g., BERT, RoBERTa, DistilBERT) to classify user inputs as either benign or prompt injection attempts. Classes Label Class Description 0 BENIGN Legitimate user queries 1 INJECTION Prompt injection attempts Features… See the full description on the dataset page: https://huggingface.co/datasets/S-Labs/prompt-injection-dataset.texttext-classification10K<n<100K8 likes742 downloads8mo agoHugging Face03yanismiraoui /prompt_injections Dataset Card for Prompt Injections by Yanis Miraoui 👋 Dataset Description This dataset of prompt injections enriches Large Language Models (LLMs) by providing task-specific examples and prompts, helping improve LLMs' performance and control their behavior. Dataset Summary This dataset contains over 1000 rows of prompt injections in multiple languages. It contains examples of prompt injections using different techniques such as: prompt leaking… See the full description on the dataset page: https://huggingface.co/datasets/yanismiraoui/prompt_injections.text1K<n<10K7 likes715 downloads4mo agoHugging Face04cgoosen /prompt_injection_password_or_secrettextn<1K3 likes270 downloads3y agoHugging Face05zachz /prompt-injection-benchmark Prompt Injection Benchmark A curated dataset of labeled prompt injection attacks and benign prompts for testing and benchmarking injection detection systems. Dataset Description This dataset contains 200 examples across 7 attack categories, plus 100 benign prompts. Each example is labeled with: text: The prompt text label: injection or benign category: Attack category (e.g., instruction_override, role_hijack) severity: low, medium, high, or critical Attack… See the full description on the dataset page: https://huggingface.co/datasets/zachz/prompt-injection-benchmark.texttext-classificationn<1K1 likes247 downloads6mo agoHugging Face06xxz224 /prompt-injection-attack-datasettabular1K<n<10K8 likes166 downloads2y agoHugging Face07cgoosen /prompt_injection_ctf_dataset_2texttext-classificationn<1K2 likes138 downloads2y agoHugging Face08MuhammadAnas1657 /Prompt_Injection_PIDStext100K<n<1M1 likes116 downloads22d agoHugging Face09rgeada /k8s-resource-prompt-injection K8s Resource Injection Dataset Dataset of real-world Kubernetes resources. Like rgeada/tool_response_injections, this dataset is constructed by randomly inserting, replacing, appending, or prepending prompt injection strings from neuralchemy/Prompt-injection-dataset into various fields of the Kubernetes resources, with the intent of training prompt-injection guardrails for agentic systems with access to Kubernetes clusters. Construction Kubernetes resource files… See the full description on the dataset page: https://huggingface.co/datasets/rgeada/k8s-resource-prompt-injection.texttext-classification10K<n<100K1 likes96 downloads1mo agoHugging Face10cgoosen /prompt_injection_combinedtabularn<1K0 likes88 downloads2y agoHugging Face11jamesdborin /Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only.tabular1K<n<10K0 likes87 downloads3mo agoHugging Face12Sahildhonde-9 /INJEXIS-Duplicate-Prompt-Injection-Dataset SPML Chatbot Prompt Injection Dataset Arxiv Paper Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sahildhonde-9/INJEXIS-Duplicate-Prompt-Injection-Dataset.tabulartext-classification10K<n<100K0 likes87 downloads23d agoHugging Face13AIDataFdn /promptinjectiontext100K<n<1M0 likes68 downloads3d agoHugging Face14hlyn-labs /prompt-injection-judge-deberta-datasetgated 🛡️ Prompt Injection Detection Dataset A 400K-sample, production-grade dataset for training binary classifiers to detect prompt injections, jailbreaks, and adversarial attacks targeting LLMs. This is the exact dataset used to train hlyn-labs/prompt-injection-judge-deberta-70m. Quick Start from datasets import load_dataset ds = load_dataset("hlyn-labs/prompt-injection-judge-deberta-dataset") Dataset Summary Stat Value Total Samples 399… See the full description on the dataset page: https://huggingface.co/datasets/hlyn-labs/prompt-injection-judge-deberta-dataset.texttext-classification100K<n<1M3 likes62 downloads4mo agoHugging Face15cgoosen /prompt_injection_ctf_dataset_3textn<1K0 likes54 downloads2y agoHugging Face16Yash0728 /Toxicity_PromptInjectiontext1K<n<10K1 likes51 downloads2y agoHugging Face17hse-llm /prompt-injectionstextn<1K0 likes48 downloads1y agoHugging Face18blackXmask /RedLockX-Prompt-Injection-109K-DataSet The RedLockX Dataset is a large-scale curated security dataset designed for evaluating and training AI systems against adversarial threats such as prompt injection, jailbreak attempts, system prompt leakage, and LLM manipulation attacks. It contains structured real-world and synthetic attack patterns used in modern AI red-teaming. 📌 Dataset Overview ✔ 109,000+ labeled adversarial & safe samples ✔ Multi-category threat… See the full description on the dataset page: https://huggingface.co/datasets/blackXmask/RedLockX-Prompt-Injection-109K-DataSet.tabulartext-classification100K<n<1M2 likes47 downloads3mo agoHugging Face19watchdogsrox /Mirror-Prompt-Injection-Dataset Mirror Prompt Injection Dataset A ~5,000-pair prompt injection detection dataset built using the Mirror design pattern, as described in: The Mirror Design Pattern: Strict Data Geometry over Model Scale for Prompt Injection Detectionhttps://arxiv.org/abs/2603.11875 Key results from the paper The paper demonstrates that a sparse character n-gram linear SVM trained on 5,000 Mirror-curated samples achieves 95.97% recall and 92.07% F1 on a holdout set, with sub-millisecond… See the full description on the dataset page: https://huggingface.co/datasets/watchdogsrox/Mirror-Prompt-Injection-Dataset.texttext-classification1K<n<10K1 likes41 downloads6mo agoHugging Face20Federico82 /prompt_injections Dataset Card for Prompt Injections by Yanis Miraoui 👋 Dataset Description This dataset of prompt injections enriches Large Language Models (LLMs) by providing task-specific examples and prompts, helping improve LLMs' performance and control their behavior. Dataset Summary This dataset contains over 1000 rows of prompt injections in multiple languages. It contains examples of prompt injections using different techniques such as: prompt leaking… See the full description on the dataset page: https://huggingface.co/datasets/Federico82/prompt_injections.text1K<n<10K0 likes32 downloads2mo agoHugging Face21Sahildhonde-9 /INJEXIS-Prompt-Injection-Dataset The RedLockX Dataset is a large-scale curated security dataset designed for evaluating and training AI systems against adversarial threats such as prompt injection, jailbreak attempts, system prompt leakage, and LLM manipulation attacks. It contains structured real-world and synthetic attack patterns used in modern AI red-teaming. 📌 Dataset Overview ✔ 109,000+ labeled adversarial & safe samples ✔ Multi-category threat… See the full description on the dataset page: https://huggingface.co/datasets/Sahildhonde-9/INJEXIS-Prompt-Injection-Dataset.tabulartext-classification100K<n<1M0 likes28 downloads2mo agoHugging Face22aestera /prompt_injection_payloadstextn<1K0 likes27 downloads2y agoHugging Face23Arthur-AI /arthur_prompt_injection_benchmarktextn<1K0 likes27 downloads1y agoHugging Face24ShieldX /Context-Aware-Repository-Prompt-Injection Overview This dataset is designed for training and evaluating AI security scanners that detect repository-aware prompt injection attacks in software development and code-assistant environments. Repository-aware prompt injections are malicious instructions embedded in code repositories, documentation, comments, configuration files, issue trackers, or other project artifacts that attempt to manipulate an AI system's behavior, override its instructions, exfiltrate sensitive… See the full description on the dataset page: https://huggingface.co/datasets/ShieldX/Context-Aware-Repository-Prompt-Injection.text1K<n<10K1 likes26 downloads3mo agoHugging Face25ClarusC64 /ai-5node-inj-buf-lag-cpl-prompt-injection-v0.1 What this repo does This dataset models prompt injection cascades in tool-using AI systems. It detects when injection pressure rises, safety buffers weaken due to incomplete filtering and trust-boundary enforcement, governance lag delays triage and revocation, and tight coupling through shared routers and scaffolds propagates injection success across products, crossing the five-node cascade threshold into an unrecoverable prompt injection cascade. This dataset models a five-node… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-inj-buf-lag-cpl-prompt-injection-v0.1.tabulartext-classificationn<1K0 likes25 downloads7mo agoHugging Face26takashi-natsume /SPML_Chatbot_Prompt_Injection SPML Chatbot Prompt Injection Dataset Arxiv Paper Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/takashi-natsume/SPML_Chatbot_Prompt_Injection.tabulartext-classification10K<n<100K1 likes20 downloads5mo agoHugging Face27shalyhinpavel /Prompt-Injection-Hard-Positivesdatasets: shalyhinpavel/Prompt-Injection-Hard-Positives The datasets (train_v3.csv and val_v3.csv) are published alongside this model atshalyhinpavel/RIG-V3-GATEKEEPER. texttext-classificationn<1K1 likes19 downloads6mo agoHugging Face28IqraParveen1 /prompt-injection-dataset Prompt Injection Detection Dataset A binary classification dataset for detecting prompt injection attacks in user inputs to LLM-based applications. Dataset Description This dataset is designed to train encoder-only models (e.g., BERT, RoBERTa, DistilBERT) to classify user inputs as either benign or prompt injection attempts. Classes Label Class Description 0 BENIGN Legitimate user queries 1 INJECTION Prompt injection attempts… See the full description on the dataset page: https://huggingface.co/datasets/IqraParveen1/prompt-injection-dataset.texttext-classification10K<n<100K0 likes18 downloads2mo agoHugging Face29mukunda1729 /prompt-injection-eval Prompt Injection Eval Set A small, hand-written evaluation set for testing prompt-injection detectors and RAG guardrails. Built as a companion to the prompt-injection-shield library. Why Most prompt-injection corpora target training (thousands of synthetic rows). This one targets evaluation: every row is hand-curated, labeled with a category and severity, and short enough to debug by eye when a detector misfires. Schema field type description text… See the full description on the dataset page: https://huggingface.co/datasets/mukunda1729/prompt-injection-eval.texttext-classificationn<1K0 likes13 downloads5mo agoHugging Face30Libertor /prompt_injections Dataset Card for Prompt Injections by Yanis Miraoui 👋 Dataset Description This dataset of prompt injections enriches Large Language Models (LLMs) by providing task-specific examples and prompts, helping improve LLMs' performance and control their behavior. Dataset Summary This dataset contains over 1000 rows of prompt injections in multiple languages. It contains examples of prompt injections using different techniques such as: prompt leaking, jailbreaking… See the full description on the dataset page: https://huggingface.co/datasets/Libertor/prompt_injections.text1K<n<10K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.