CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01deepset /prompt-injections Dataset Card for "deberta-v3-base-injection-dataset" More Information needed textn<1K183 likes11k downloads2y agoHugging Face02neuralchemy /Prompt-injection-dataset advance dataset if you want for llm security https://huggingface.co/datasets/neuralchemy/prompt-injection-Threat-Matrix Prompt Injection & Jailbreak Detection Dataset A high-quality, leakage-free binary classification dataset for detecting prompt injection and jailbreak attacks against Large Language Models. Zero data leakage — group-aware splitting confirmed Balanced classes — ~60% malicious / 40% benign Two configs — core for classical ML, full for transformers 29… See the full description on the dataset page: https://huggingface.co/datasets/neuralchemy/Prompt-injection-dataset.texttext-classification10K<n<100K30 likes2.8k downloads5mo agoHugging Face03rogue-security /prompt-injections-benchmarkgated Dataset: Qualifire Benchmark Prompt Injection(Jailbreak vs. Benign) Datasets Overview This dataset contains 5,000 prompts, each labeled as either jailbreak or benign. The dataset is designed for evaluating AI models' robustness against adversarial prompts and their ability to distinguish between safe and unsafe inputs. Dataset Structure Total Samples: 5,000 Labels: jailbreak, benign Columns: text: The input text label: The classification (jailbreak or benign)… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/prompt-injections-benchmark.text1K<n<10K45 likes1k downloads6mo agoHugging Face04jayavibhav /prompt-injection-safetytext10K<n<100K13 likes949 downloads2y agoHugging Face05S-Labs /prompt-injection-dataset Prompt Injection Detection Dataset A binary classification dataset for detecting prompt injection attacks in user inputs to LLM-based applications. Dataset Description This dataset is designed to train encoder-only models (e.g., BERT, RoBERTa, DistilBERT) to classify user inputs as either benign or prompt injection attempts. Classes Label Class Description 0 BENIGN Legitimate user queries 1 INJECTION Prompt injection attempts Features… See the full description on the dataset page: https://huggingface.co/datasets/S-Labs/prompt-injection-dataset.texttext-classification10K<n<100K8 likes767 downloads8mo agoHugging Face06yanismiraoui /prompt_injections Dataset Card for Prompt Injections by Yanis Miraoui 👋 Dataset Description This dataset of prompt injections enriches Large Language Models (LLMs) by providing task-specific examples and prompts, helping improve LLMs' performance and control their behavior. Dataset Summary This dataset contains over 1000 rows of prompt injections in multiple languages. It contains examples of prompt injections using different techniques such as: prompt leaking… See the full description on the dataset page: https://huggingface.co/datasets/yanismiraoui/prompt_injections.text1K<n<10K7 likes734 downloads4mo agoHugging Face07jayavibhav /prompt-injectiontext100K<n<1M8 likes691 downloads2y agoHugging Face08JasperLS /prompt-injections Dataset Card for "deberta-v3-base-injection-dataset" More Information needed textn<1K21 likes440 downloads3y agoHugging Face09wambosec /prompt-injections wambosec/prompt-injections A dataset of prompts for training prompt injection detection models. Dataset Description This dataset contains prompts labeled as either benign (normal user requests) or malicious (prompt injection attacks). Dataset Statistics Total prompts: 5,766 Benign prompts: 2,340 Malicious prompts: 3,426 Malicious ratio: 59.4% Dataset Structure { "prompt": str, # The prompt text "label": int, # 0 =… See the full description on the dataset page: https://huggingface.co/datasets/wambosec/prompt-injections.texttext-classification1K<n<10K3 likes346 downloads8mo agoHugging Face10imoxto /prompt_injection_cleaned_dataset-v2 Dataset Card for "prompt_injection_cleaned_dataset-v2" More Information needed text100K<n<1M11 likes308 downloads3y agoHugging Face11neuralchemy /prompt-injection-Threat-Matrix CATEGORIZED DATASET - easy to use https://huggingface.co/datasets/neuralchemy/prompt-injection-dataset-categorized Neuralchemy Prompt Injection Threat Matrix A professional-grade prompt injection and jailbreak detection dataset featuring 32,320 curated samples across 5 dimensions with full threat intelligence schema including technique classification, severity scoring, attack surface detection, and ambiguity flagging. Built for training production-grade LLM… See the full description on the dataset page: https://huggingface.co/datasets/neuralchemy/prompt-injection-Threat-Matrix.tabulartext-classification10K<n<100K3 likes286 downloads3mo agoHugging Face12cgoosen /prompt_injection_password_or_secrettextn<1K3 likes267 downloads3y agoHugging Face13hirundo-io /prompt-injection-purple-llamatextn<1K0 likes251 downloads1y agoHugging Face14neuralchemy /prompt-injection-dataset-categorized Prompt Injection Dataset — Categorized (Threat Matrix V2) Welcome to Prompt Injection Dataset – Categorized (formerly Threat Matrix), by Neuralchemy. This is the successor to our original Prompt Injection Threat Matrix dataset. Instead of one multi-label table, this version splits the taxonomy into 7 clean, single-purpose subsets — 6 taxonomy dimensions plus a bonus ambiguity flag — so you can train a focused specialist model on each one instead of fighting multi-task learning.… See the full description on the dataset page: https://huggingface.co/datasets/neuralchemy/prompt-injection-dataset-categorized.tabulartext-classification100K<n<1M1 likes251 downloads2mo agoHugging Face15imoxto /prompt_injection_cleaned_dataset Dataset Card for "prompt_injection_cleaned_dataset" More Information needed tabular100K<n<1M6 likes228 downloads3y agoHugging Face16zachz /prompt-injection-benchmark Prompt Injection Benchmark A curated dataset of labeled prompt injection attacks and benign prompts for testing and benchmarking injection detection systems. Dataset Description This dataset contains 200 examples across 7 attack categories, plus 100 benign prompts. Each example is labeled with: text: The prompt text label: injection or benign category: Attack category (e.g., instruction_override, role_hijack) severity: low, medium, high, or critical Attack… See the full description on the dataset page: https://huggingface.co/datasets/zachz/prompt-injection-benchmark.texttext-classificationn<1K1 likes228 downloads6mo agoHugging Face17prodnull /prompt-injection-repo-datasetgated Prompt Injection Repository File Dataset A labeled dataset for detecting prompt injection attacks in repository files — code, configs, READMEs, CI/CD workflows, and documentation that AI coding agents process as context. What This Is (and Isn't) This dataset targets a specific threat: indirect prompt injection via repository content. When AI coding agents (Claude Code, Cursor, Copilot, Gemini CLI) clone a repo, every file becomes part of the agent's context.… See the full description on the dataset page: https://huggingface.co/datasets/prodnull/prompt-injection-repo-dataset.texttext-classification1K<n<10K11 likes194 downloads7mo agoHugging Face18geekyrakshit /prompt-injection-datasetCollected from the following datasets: deepset/prompt-injections xTRam1/safe-guard-prompt-injection jayavibhav/prompt-injection text100K<n<1M8 likes180 downloads2y agoHugging Face19xxz224 /prompt-injection-attack-datasettabular1K<n<10K8 likes171 downloads2y agoHugging Face20ivanleomk /prompt_injection_password Dataset Card for "prompt_injection_password" More Information needed textn<1K1 likes161 downloads3y agoHugging Face21rikka-snow /prompt-injection-multilingualtexttext-classification1K<n<10K1 likes152 downloads2y agoHugging Face22issdandavis /prompt-injection-bit-signatures Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data. Prompt Injection → Bit Signatures 24,254 labeled prompts from 4 public prompt-injection datasets, each mapped through the Six Sacred Tongues bijective tokenizer from the SCBE-AETHERMOORE framework into a lossless per-prompt bit signature. Stratified 70/15/15 train/val/test split by (source, label) so every source is represented in every split with its original label… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/prompt-injection-bit-signatures.texttext-classification10K<n<100K2 likes152 downloads2mo agoHugging Face23cgoosen /prompt_injection_ctf_dataset_2texttext-classificationn<1K2 likes140 downloads2y agoHugging Face24mangalathkedar /prompt-injection-multilayertext10K<n<100K0 likes128 downloads3mo agoHugging Face25cyberec /Prompt-injection-dataset Prompt Injection & Jailbreak Detection Dataset A high-quality, leakage-free binary classification dataset for detecting prompt injection and jailbreak attacks against Large Language Models. Zero data leakage — group-aware splitting confirmed Balanced classes — ~60% malicious / 40% benign Two configs — core for classical ML, full for transformers 29 attack categories including cutting-edge 2025 techniques Severity labels, source tracking, augmentation flags on every row… See the full description on the dataset page: https://huggingface.co/datasets/cyberec/Prompt-injection-dataset.texttext-classification10K<n<100K2 likes126 downloads5mo agoHugging Face26v1adam /Prompt_injection_and_Sensitive_Data_exposure_detectiontext1K<n<10K2 likes125 downloads2mo agoHugging Face27aporia-ai /prompt_injectiontextn<1K1 likes123 downloads2y agoHugging Face28MuhammadAnas1657 /Prompt_Injection_PIDStext100K<n<1M1 likes116 downloads22d agoHugging Face29anggiatm /prompt-injection-gemmatext1K<n<10K0 likes105 downloads2y agoHugging Face30Alignment-Lab-AI /Prompt-Injection-Testtext1K<n<10K0 likes103 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.