CoolFace
Datasetpublic

rupeshreddypapa/nepi-prompts-dataset

NEPI: Narrative-Embedded Prompt Injection Dataset (Sanitized) Dataset Summary This dataset contains 4,000 sanitized prompts designed for research on prompt injection vulnerabilities in Large Language Models (LLMs).It introduces and supports evaluation of a novel attack class called Narrative-Embedded Prompt Injection (NEPI), where adversarial intent is embedded inside coherent fictional narratives, dialogues, or persona-driven roleplay prompts. Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/rupeshreddypapa/nepi-prompts-dataset.

sourceHugging Facemitupdated 8d agoView on Hugging Face
0likes38downloads
Dataset Card

NEPI: Narrative-Embedded Prompt Injection Dataset (Sanitized)

Dataset Summary

This dataset contains 4,000 sanitized prompts designed for research on prompt injection vulnerabilities in Large Language Models (LLMs). It introduces and supports evaluation of a novel attack class called Narrative-Embedded Prompt Injection (NEPI), where adversarial intent is embedded inside coherent fictional narratives, dialogues, or persona-driven roleplay prompts.

Unlike traditional jailbreak prompts that explicitly override system instructions, NEPI prompts exploit the model’s tendency to maintain narrative coherence and follow implicit role-based context.

This dataset is intended for academic research, benchmarking, and defensive model training.


Motivation

Prompt injection remains a major unresolved challenge for instruction-tuned LLMs. Existing defenses (keyword filters, boundary enforcement, and prompt sanitization) often fail against narrative-based attacks, because malicious intent is not presented as explicit commands but instead as part of a creative narrative flow.

This dataset was created as part of the research project:

"Narrative-Embedded Prompt Injection in Large Language Models: Attack Characterization and Defense Strategies"


Dataset Composition

The dataset contains a balanced mix of benign prompts, traditional jailbreak attacks, and NEPI attack variants.

Total samples: 4000

CategoryAttack FamilyLabelCount
Benign QABenignSafe250
Benign Creative WritingBenignSafe250
Traditional JailbreaksTraditionalMalicious300+
Instruction OverrideTraditionalMalicious250+
Delimiter InjectionTraditionalMalicious250+
Role-based NEPINEPIMalicious700+
Dialogue-based NEPINEPIMalicious700+
Monologue-based NEPINEPIMalicious600+
Mixed-Intent NEPINEPIMalicious500+
Ambiguous Narrative PromptsNEPISuspicious200

Labels

Prompts are annotated into three categories:

  • —Safe: benign prompts with no malicious intent.
  • —Suspicious: ambiguous narrative prompts that may contain hidden intent or role manipulation but do not explicitly enforce adversarial behavior.
  • —Malicious: prompt injection attempts designed to manipulate model behavior (traditional jailbreaks or NEPI variants).

Dataset Features

Each sample contains the following fields:

  • —id: unique identifier for the prompt.
  • —prompt: the sanitized prompt text.
  • —label: one of {Safe, Suspicious, Malicious}.
  • —attack_family: one of {Benign, Traditional, NEPI}.
  • —attack_type: specific category (e.g., role_based, dialogue_based, traditional_jailbreak).
  • —target_goal: intended adversarial objective category.
  • —source: dataset source (synthetic_template).

Sanitization and Responsible Disclosure

This dataset is sanitized for public release.

The dataset preserves the structure and semantics of narrative prompt injection attacks but replaces explicit harmful or operational payloads with neutral placeholders such as:

  • —[SENSITIVE_INFO]
  • —[FINANCIAL_INFO]
  • —[REDACTED_OBJECT]

This ensures the dataset can be safely used for defense research, benchmarking, and classifier training without enabling real-world misuse.


Intended Uses

Recommended Uses

  • —Benchmarking LLM prompt injection robustness
  • —Training and evaluating prompt-injection classifiers
  • —Studying narrative-based jailbreak vulnerabilities
  • —Red-teaming research under controlled settings
  • —Defense pipeline evaluation (input filtering + output validation)

Not Recommended Uses

  • —Deploying prompts as-is in real systems
  • —Attempting to bypass safety systems in production environments

Citation

If you use this dataset in your research, please cite:

bibtex
@article{pawar2026nepi,
  title={Narrative-Embedded Prompt Injection in Large Language Models: Attack Characterization and Defense Strategies},
  author={Pawar, Vaibhav and Urvashi},
  journal={Under Review},
  year={2026}
}
License
This dataset is released under the MIT License.

Authors
Vaibhav Pawar, Dr B R Ambedkar NIT Jalandhar

Dr Urvashi, Dr B R Ambedkar NIT Jalandhar

Links
Paper Code Repository (GitHub): https://github.com/Vaibhav-GOAT/NEPI-Attack-Defense

Contact
For questions or collaborations:

📧 vaibhavpp.is.24@nitj.ac.in
📧 vaibhavpitambarpawar69@gmail.com