datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
prompts.chat
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can… See the full description on the dataset page: https://huggingface.co/datasets/fka/prompts.chat.real-toxicity-prompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/allenai/real-toxicity-prompts.mt_bench_prompts
MT Bench by LMSYS
This set of evaluation prompts is created by the LMSYS org for better evaluation of chat models.
For more information, see the paper.
Dataset loading
To load this dataset, use 🤗 datasets:
from datasets import load_dataset
data = load_dataset(HuggingFaceH4/mt_bench_prompts, split="train")
Dataset creation
To create the dataset, we do the following for our internal tooling.
rename turns to prompts,
add empty reference to remaining prompts… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/mt_bench_prompts.rl-lm-formality-promptsChatGPT-Jailbreak-Promptsrl-lm-imdb-promptstm-system_promptdrh-System-Prompt-processedprompt-injections
Dataset Card for "deberta-v3-base-injection-dataset"
More Information needed
rl-lm-toxicity-promptsPromptEval_MMLU_full
MMLU Multi-Prompt Evaluation Data
Overview
This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in
Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. "Efficient multi-prompt evaluation of LLMs."… See the full description on the dataset page: https://huggingface.co/datasets/PromptEval/PromptEval_MMLU_full.verbalist_prompts
Verbalist (буквоед) - русскоязычный ассистент.
Проект во многом вдохновленный Saiga.
Мною были собраны все самые качественные датасеты с huggingface.datasets, а также собраны дополнительно с тех сайтов, которые я посчитал весьма полезными для создания аналога ChatGPT. Лицензии у всех датасетов отличаются, какие-то по типу OpenAssistant/oasst1 были созданы специально для обучения подобных моделей, какие-то являются прямой выгрузкой диалогов с ChatGPT (RyokoAI/ShareGPT52K).
Вклад… See the full description on the dataset page: https://huggingface.co/datasets/DeepPavlov/verbalist_prompts.hle_prompts_07_02_25in-the-wild-jailbreak-prompts
In-The-Wild Jailbreak Prompts on LLMs
This is the official repository for the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models by Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang.
In this project, employing our new framework JailbreakHub, we conduct the first measurement study on jailbreak prompts in the wild, with 15,140 prompts collected from December 2022 to December 2023 (including 1,405… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/in-the-wild-jailbreak-prompts.ChatGPT-4o-Writing-Prompts
ChatGPT-4o Writing Prompts
This is a dataset containing 3746 short stories, generated with OpenAI's chatgpt-4o-latest model and using Reddit's Writing Prompts subreddit as a source. Each sample is generally between 6000-8000 characters long.
These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres.
Note that I did not touch the Markdown ChatGPT-4o produced by itself to enrich its output, as I very much enjoy the added flavour… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts.open-models-prompt-datasets
🖼️ Open Models Prompt Dataset
🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.prompt-perfect
Scoring popular datasets with "Self-Alignment with Instruction Backtranslation" prompt
35 datasets scored (>6B tokens)
Scoring Models used
gpt-3.5-turbo-16k
gpt-3.5-turbo-1106
gpt-3.5-turbo-0125
All datasets have 2 additional columns
score - Response from the model including CoT (if provided)
extracted_score - Extracted score from the score column as int
Datasets Scored by Prompt (Needs to be updated)… See the full description on the dataset page: https://huggingface.co/datasets/0-hero/prompt-perfect.messy-prompt-datasets
🎨 Messy Prompt Dataset
🎨 A mixed collection of AI image prompts (500+). A bit of everything — raw and uncurated. Truly open source: No login, no ads, no redirection. Just pure data for AI creators.
This project is a growing collection of diverse image generation prompts gathered from social platforms like Twitter/X. The entire dataset contains 500+ images, all structured into a comprehensive dataset.
Due to GitHub's limitations with large file storage, the full dataset… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/messy-prompt-datasets.Stable-Diffusion-Prompts
Stable Diffusion Dataset
This is a set of about 80,000 prompts filtered and extracted from the image finder for Stable Diffusion: "Lexica.art". It was a little difficult to extract the data, since the search engine still doesn't have a public API without being protected by cloudflare.
If you want to test the model with a demo, you can go to: "spaces/Gustavosta/MagicPrompt-Stable-Diffusion".
If you want to see the model, go to: "Gustavosta/MagicPrompt-Stable-Diffusion".
veo3-video-prompts
Veo 3 Video Generation Dataset
English | Português do Brasil
English
Summary
A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant.
Videos: 5,811
Input images: 1,354
Configurations: 6
Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.AIME_Problem_Set_1983-2024Prompt-injection-dataset
advance dataset if you want for llm security
https://huggingface.co/datasets/neuralchemy/prompt-injection-Threat-Matrix
Prompt Injection & Jailbreak Detection Dataset
A high-quality, leakage-free binary classification dataset for detecting prompt injection and jailbreak attacks against Large Language Models.
Zero data leakage — group-aware splitting confirmed
Balanced classes — ~60% malicious / 40% benign
Two configs — core for classical ML, full for transformers
29… See the full description on the dataset page: https://huggingface.co/datasets/neuralchemy/Prompt-injection-dataset.parti-prompts
Dataset Card for PartiPrompts (P2)
Dataset Summary
PartiPrompts (P2) is a rich set of over 1600 prompts in English that we release
as part of this work. P2 can be used to measure model capabilities across
various categories and challenge aspects.
P2 prompts can be simple, allowing us to gauge the progress from scaling. They
can also be complex, such as the following 67-word description we created for
Vincent van Gogh’s The Starry Night (1889):
Oil-on-canvas painting of a… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/parti-prompts.cyberseceval3-visual-prompt-injection
Dataset Card for CyberSecEval 3 - Visual Prompt Injection Benchmark
Dataset Details
Dataset Description
This dataset provides a multimodal benchmark for visual prompt injection, with text/image inputs. It is part of CyberSecEval 3, the third edition of Meta's flagship suite of security benchmarks for LLMs to measure cybersecurity risks and capabilities across multiple domains.
Language(s): English
License: MIT
Dataset Sources
Repository: Link… See the full description on the dataset page: https://huggingface.co/datasets/facebook/cyberseceval3-visual-prompt-injection.product-photography-v1-tiny-prompts-tasks-collage-filteredsafe-guard-prompt-injectionWe formulated the prompt injection detector problem as a classification problem and trained our own language model
to detect whether a given user prompt is an attack or safe. First, to train our own prompt injection detector, we
required high-quality labelled data; however, existing prompt injection datasets were either too small (on the magnitude
of O(100)) or didn’t cover a broad spectrum of prompt injection attacks. To this end, inspired by the GLAN paper,
we created a custom synthetic… See the full description on the dataset page: https://huggingface.co/datasets/xTRam1/safe-guard-prompt-injection.ioi-eval-openrouter_anthropic_claude-3_7-sonnet_thinking-prompt-mem-limitioi-eval-openrouter_openai_gpt-3.5-turbo-new-promptioi-eval-openrouter_openai_gpt-3.5-turbo-prompt-mem-limit
