datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
prompts.chat
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can… See the full description on the dataset page: https://huggingface.co/datasets/fka/prompts.chat.hle_prompts_07_02_25parti-prompts
Dataset Card for PartiPrompts (P2)
Dataset Summary
PartiPrompts (P2) is a rich set of over 1600 prompts in English that we release
as part of this work. P2 can be used to measure model capabilities across
various categories and challenge aspects.
P2 prompts can be simple, allowing us to gauge the progress from scaling. They
can also be complex, such as the following 67-word description we created for
Vincent van Gogh’s The Starry Night (1889):
Oil-on-canvas painting of a… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/parti-prompts.SPML_Chatbot_Prompt_Injection
SPML Chatbot Prompt Injection Dataset
Arxiv Paper
Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/reshabhs/SPML_Chatbot_Prompt_Injection.prompt-injection-dataset
Prompt Injection Detection Dataset
A binary classification dataset for detecting prompt injection attacks in user inputs to LLM-based applications.
Dataset Description
This dataset is designed to train encoder-only models (e.g., BERT, RoBERTa, DistilBERT) to classify user inputs as either benign or prompt injection attempts.
Classes
Label
Class
Description
0
BENIGN
Legitimate user queries
1
INJECTION
Prompt injection attempts
Features… See the full description on the dataset page: https://huggingface.co/datasets/S-Labs/prompt-injection-dataset.prompt_injections
Dataset Card for Prompt Injections by Yanis Miraoui 👋
Dataset Description
This dataset of prompt injections enriches Large Language Models (LLMs) by providing task-specific examples and prompts, helping improve LLMs' performance and control their behavior.
Dataset Summary
This dataset contains over 1000 rows of prompt injections in multiple languages. It contains examples of prompt injections using different techniques such as: prompt leaking… See the full description on the dataset page: https://huggingface.co/datasets/yanismiraoui/prompt_injections.Nemotron-SFT-Agentic-v2-prompt-only
Nemotron-SFT-Agentic-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Agentic-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Agentic-v2-prompt-only.llm-cost-same-prompt
Measured per-call LLM cost — same prompt, every model
Vendors publish prices per million tokens. Nobody publishes what one call actually costs, because
that depends on how many tokens the model chooses to emit — and on the same question models differ by
more than an order of magnitude. One model finishes a JSON extraction in 23 tokens; another writes 300.
This dataset sends a fixed set of prompts to every model at temperature 0, every night, and records
the cost computed from… See the full description on the dataset page: https://huggingface.co/datasets/mario0369/llm-cost-same-prompt.text-to-image-prompts
The dataset of the most popular text-to-image prompts.
Dataset Details
Dataset Description
Curated by: kazimir.ai
Funded by [optional]: [More Information Needed]
Shared by [optional]: https://kazimir.ai
License: apache-2.0
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Free to use.
Dataset Structure
CSV file… See the full description on the dataset page: https://huggingface.co/datasets/Kazimir-ai/text-to-image-prompts.CCP-sensitive-prompts
CCP Sensitive Prompts
These prompts cover sensitive topics in China, and are likely to be censored by Chinese models.
prompt_injection_password_or_secretNemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only.prompt-injection-benchmark
Prompt Injection Benchmark
A curated dataset of labeled prompt injection attacks and benign prompts for testing and benchmarking injection detection systems.
Dataset Description
This dataset contains 200 examples across 7 attack categories, plus 100 benign prompts. Each example is labeled with:
text: The prompt text
label: injection or benign
category: Attack category (e.g., instruction_override, role_hijack)
severity: low, medium, high, or critical
Attack… See the full description on the dataset page: https://huggingface.co/datasets/zachz/prompt-injection-benchmark.ChatGPT-prompts
ChatGPT-Prompts Dataset
Description
This dataset aims to provide an evaluation data for the Language Models to come. It has been generated using LearnGPT website.
prompt-injection-attack-datasetAssistive_Prompting_Disabilities_Dataset
Assistive Prompting Disabilities Dataset
This dataset provides multilingual prompts for evaluating assistive prompt mediation under accessibility-related textual noise. It accompanies the ICML accepted Assistive Prompt Mediation paper and includes benchmark scripts for preparing inference inputs, computing row-level metrics, compiling existing judge annotations, and generating aggregate summaries.
Dataset Description
The dataset contains clean prompts and noisy… See the full description on the dataset page: https://huggingface.co/datasets/AssistivePromptMediation/Assistive_Prompting_Disabilities_Dataset.videophy2_upsampled_promptsprompt_injection_ctf_dataset_2PromptEvalsPromptEvals: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines
Large language models (LLMs) are increasingly deployed in specialized production data processing pipelines across diverse domains---such as finance, marketing, and e-commerce.
However, when running them in production across many inputs, they often fail to follow instructions or meet developer expectations.
To improve reliability in these applications, creating assertions or guardrails… See the full description on the dataset page: https://huggingface.co/datasets/reyavir/PromptEvals.marketing_user_prompts_unfilteredPromptCloudHQ_flipkart-products
Flipkart Products
20,000 products on Flipkart
Dataset Info
Source: Kaggle
Original Size: 5.50 MB
Kaggle Downloads: 29,395
Files: 1
Files
flipkart_com-ecommerce_sample.csv
Mirrored from Kaggle
10k_rows_cleaned_prompts
10K Rows Cleaned Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
You must provide attribution when using this data in publications, research, or commercial products.
Dataset Overview
A chunked collection of 2.7 million cleaned English prompts, organized into 200 files of 10,000 rows each for easy processing and distributed training of language models.
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/10k_rows_cleaned_prompts.synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.Prompt_Injection_PIDSmalicious-promptsAI-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
Jailbreak Prompts
Dataset Summary
Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[German]
prompttest-18videosprompts_under_512_tokens
Under 512 Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
Dataset Overview
Specialized collection of short-form English prompts (under 512 tokens), perfect for training models with context length constraints or faster iteration cycles.
📊 Dataset Statistics
Metric
Value
Total Files
200
Rows Per File
10,000
Total Rows
2,000,000
Token Range
1 to 511 tokens… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/prompts_under_512_tokens.malicious-prompts-minilm-embeddings
