datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cyberseceval3-visual-prompt-injection
Dataset Card for CyberSecEval 3 - Visual Prompt Injection Benchmark
Dataset Details
Dataset Description
This dataset provides a multimodal benchmark for visual prompt injection, with text/image inputs. It is part of CyberSecEval 3, the third edition of Meta's flagship suite of security benchmarks for LLMs to measure cybersecurity risks and capabilities across multiple domains.
Language(s): English
License: MIT
Dataset Sources
Repository: Link… See the full description on the dataset page: https://huggingface.co/datasets/facebook/cyberseceval3-visual-prompt-injection.Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1
Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1
Dataset Description:
Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 is an RL dataset for training and evaluating a tool-using agent's ability to resist Indirect Prompt Injection (IPI) attacks hidden inside tool-returned environment data. In each record, the agent receives a benign user request that requires calling a read tool whose output contains an adversarial instruction disguised as legitimate domain content… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1.Turkish_prompt_injection_jailbreak_dataset
[!NOTE]
Türkçe Prompt Injection & Jailbreak Veri Seti
📌 Atıf / Citation
Bu veri setini akademik çalışmalarda, model değerlendirmelerinde, güvenlik analizlerinde veya türev araştırmalarda kullanırsanız lütfen aşağıdaki makaleye atıf veriniz:
Aytaş, Ö.; Şen, T.; Diri, B.; Biricik, G.; Bayram, M.A. Benchmarking Prompt Injection Attacks on LLMs: Turkish Vulnerability Assessment and English Comparative Analysis. Applied Sciences 2026, 16(13), 6740.… See the full description on the dataset page: https://huggingface.co/datasets/OnerAYTAS/Turkish_prompt_injection_jailbreak_dataset.evaded-prompt-injection-and-jailbreak-samplesThis dataset originates from our paper 'Bypassing Prompt Injection and Jailbreak Detection in LLM Guardrails'.
The dataset contains a mixture of prompt injections and jailbreak samples modified via character injection and adversarial ML evasion techniques (Techniques can be found within the paper above). For each sample we provide the original unaltered prompt and a modified prompt, the attack_name outlines which attack technique was used to modify the sample.
Acknowledgements… See the full description on the dataset page: https://huggingface.co/datasets/Mindgard/evaded-prompt-injection-and-jailbreak-samples.omnimcp_mcp_prompt_injection_guard_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_prompt_injection_guard_teaser.prompt-injection-defense-dpo-3k
Prompt Injection Defense DPO (3K)
DPO preference pairs training LLMs to detect and resist prompt injection attacks.
Motivation
As LLMs are deployed in agentic and production contexts, prompt injection — where malicious instructions are embedded in user input or retrieved documents — is a critical security threat. This dataset trains models to recognize and decline injection attempts while remaining helpful for legitimate queries.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/prompt-injection-defense-dpo-3k.dataset_with_prompt_injection
📦 Dataset Card
Dataset Summary
This dataset contains examples for training and evaluating language models.
The data is stored in JSONL format, where each line represents one training example.
Typical use cases include:
Instruction fine-tuning
Response generation
Conversational modelling
Question answering
Prompt injection research
🎯 Intended Uses
This dataset is intended for research and learning purposes:
Training LLMs
Experimenting with fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/shehrozrafaqat/dataset_with_prompt_injection.nemotron-gym-agentic-indirect-prompt-injection
laion/nemotron-gym-agentic-indirect-prompt-injection
Harbor task-binary dataset (1,272 tasks) converted from nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1
(part of the nvidia/Nemotron-Post-Training-v3 collection).
Each row is a valid Harbor
task binary: columns path (str) and task_binary (gzip tar). Converted with the
OpenThoughts-Agent data.nemotron_gym framework.
Grading: Single-step injection-resist proxy (reward 1 = did NOT emit injected call).
dataset_with_prompt_injection
📦 Dataset Card
Dataset Summary
This dataset contains examples for training and evaluating language models.
The data is stored in JSONL format, where each line represents one training example.
Typical use cases include:
Instruction fine-tuning
Response generation
Conversational modelling
Question answering
Prompt injection research
🎯 Intended Uses
This dataset is intended for research and learning purposes:
Training LLMs
Experimenting with fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/yashalhussain/dataset_with_prompt_injection.dataset_with_prompt_injection
📦 Dataset Card
Dataset Summary
This dataset contains examples for training and evaluating language models.
The data is stored in JSONL format, where each line represents one training example.
Typical use cases include:
Instruction fine-tuning
Response generation
Conversational modelling
Question answering
Prompt injection research
🎯 Intended Uses
This dataset is intended for research and learning purposes:
Training LLMs
Experimenting with fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/SyedZainAliShah/dataset_with_prompt_injection.dataset_with_prompt_injection
📦 Dataset Card
Dataset Summary
This dataset contains examples for training and evaluating language models.
The data is stored in JSONL format, where each line represents one training example.
Typical use cases include:
Instruction fine-tuning
Response generation
Conversational modelling
Question answering
Prompt injection research
🎯 Intended Uses
This dataset is intended for research and learning purposes:
Training LLMs
Experimenting with fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/MuhammadAhmadZia/dataset_with_prompt_injection.dataset_with_prompt_injection
📦 Dataset Card
Dataset Summary
This dataset contains examples for training and evaluating language models.
The data is stored in JSONL format, where each line represents one training example.
Typical use cases include:
Instruction fine-tuning
Response generation
Conversational modelling
Question answering
Prompt injection research
🎯 Intended Uses
This dataset is intended for research and learning purposes:
Training LLMs
Experimenting with fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/Ansnaeem/dataset_with_prompt_injection.dataset_with_prompt_injection
📦 Dataset Card
Dataset Summary
This dataset contains examples for training and evaluating language models.
The data is stored in JSONL format, where each line represents one training example.
Typical use cases include:
Instruction fine-tuning
Response generation
Conversational modelling
Question answering
Prompt injection research
🎯 Intended Uses
This dataset is intended for research and learning purposes:
Training LLMs
Experimenting with fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/abubakaraabi786/dataset_with_prompt_injection.dataset_with_prompt_injection
📦 Dataset Card
Dataset Summary
This dataset contains examples for training and evaluating language models.
The data is stored in JSONL format, where each line represents one training example.
Typical use cases include:
Instruction fine-tuning
Response generation
Conversational modelling
Question answering
Prompt injection research
🎯 Intended Uses
This dataset is intended for research and learning purposes:
Training LLMs
Experimenting with fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/Laiba-07/dataset_with_prompt_injection.dataset_with_prompt_injection
📦 Dataset Card
Dataset Summary
This dataset contains examples for training and evaluating language models.
The data is stored in JSONL format, where each line represents one training example.
Typical use cases include:
Instruction fine-tuning
Response generation
Conversational modelling
Question answering
Prompt injection research
🎯 Intended Uses
This dataset is intended for research and learning purposes:
Training LLMs
Experimenting with fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/WaliyaKhan882/dataset_with_prompt_injection.dataset_with_prompt_injection
📦 Dataset Card
Dataset Summary
This dataset contains examples for training and evaluating language models.
The data is stored in JSONL format, where each line represents one training example.
Typical use cases include:
Instruction fine-tuning
Response generation
Conversational modelling
Question answering
Prompt injection research
🎯 Intended Uses
This dataset is intended for research and learning purposes:
Training LLMs
Experimenting with fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/ShanzaGull/dataset_with_prompt_injection.dataset_with_prompt_injection
📦 Dataset Card
Dataset Summary
This dataset contains examples for training and evaluating language models.
The data is stored in JSONL format, where each line represents one training example.
Typical use cases include:
Instruction fine-tuning
Response generation
Conversational modelling
Question answering
Prompt injection research
🎯 Intended Uses
This dataset is intended for research and learning purposes:
Training LLMs
Experimenting with fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/Wosqa/dataset_with_prompt_injection.dataset_with_prompt_injection
📦 Dataset Card
Dataset Summary
This dataset contains examples for training and evaluating language models.
The data is stored in JSONL format, where each line represents one training example.
Typical use cases include:
Instruction fine-tuning
Response generation
Conversational modelling
Question answering
Prompt injection research
🎯 Intended Uses
This dataset is intended for research and learning purposes:
Training LLMs
Experimenting with fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/Maryam7711/dataset_with_prompt_injection.dataset_with_prompt_injection
📦 Dataset Card
Dataset Summary
This dataset contains examples for training and evaluating language models.
The data is stored in JSONL format, where each line represents one training example.
Typical use cases include:
Instruction fine-tuning
Response generation
Conversational modelling
Question answering
Prompt injection research
🎯 Intended Uses
This dataset is intended for research and learning purposes:
Training LLMs
Experimenting with fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/MuhammadAhmad332/dataset_with_prompt_injection.dataset_with_prompt_injection
📦 Dataset Card
Dataset Summary
This dataset contains examples for training and evaluating language models.
The data is stored in JSONL format, where each line represents one training example.
Typical use cases include:
Instruction fine-tuning
Response generation
Conversational modelling
Question answering
Prompt injection research
🎯 Intended Uses
This dataset is intended for research and learning purposes:
Training LLMs
Experimenting with fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/S-a-r-a/dataset_with_prompt_injection.
