CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /real-toxicity-prompts Dataset Card for Real Toxicity Prompts Dataset Summary RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models. Languages English Dataset Structure Data Instances Each instance represents a prompt and its metadata: { "filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt", "begin":340, "end":564, "challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/allenai/real-toxicity-prompts.tabular10K<n<100K123 likes20k downloads4y agoHugging Face02garak-llm /drh-System-Prompt-processedtextn<1K0 likes13k downloads5mo agoHugging Face03garak-llm /tm-system_prompttextn<1K0 likes13k downloads8mo agoHugging Face04Gryphe /ChatGPT-4o-Writing-Prompts ChatGPT-4o Writing Prompts This is a dataset containing 3746 short stories, generated with OpenAI's chatgpt-4o-latest model and using Reddit's Writing Prompts subreddit as a source. Each sample is generally between 6000-8000 characters long. These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres. Note that I did not touch the Markdown ChatGPT-4o produced by itself to enrich its output, as I very much enjoy the added flavour… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts.texttext-generation1K<n<10K36 likes6.1k downloads2y agoHugging Face05artificialguybr /veo3-video-prompts Veo 3 Video Generation Dataset English | Português do Brasil English Summary A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant. Videos: 5,811 Input images: 1,354 Configurations: 6 Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.imagetext-to-video1K<n<10K0 likes5.3k downloads1mo agoHugging Face06facebook /cyberseceval3-visual-prompt-injection Dataset Card for CyberSecEval 3 - Visual Prompt Injection Benchmark Dataset Details Dataset Description This dataset provides a multimodal benchmark for visual prompt injection, with text/image inputs. It is part of CyberSecEval 3, the third edition of Meta's flagship suite of security benchmarks for LLMs to measure cybersecurity risks and capabilities across multiple domains. Language(s): English License: MIT Dataset Sources Repository: Link… See the full description on the dataset page: https://huggingface.co/datasets/facebook/cyberseceval3-visual-prompt-injection.imagetext-generation1K<n<10K10 likes2.7k downloads2y agoHugging Face07nvidia /Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 Dataset Description: Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 is an RL dataset for training and evaluating a tool-using agent's ability to resist Indirect Prompt Injection (IPI) attacks hidden inside tool-returned environment data. In each record, the agent receives a benign user request that requires calling a read tool whose output contains an adversarial instruction disguised as legitimate domain content… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1.textreinforcement-learning1K<n<10K8 likes1.4k downloads4mo agoHugging Face08xl-zhao /PromptCoT-2.0-SFT-4.8M PromptCoT-2.0-SFT-4.8M This repository contains the largest dataset released with PromptCoT 2.0 (Scaling Prompt Synthesis for LLM Reasoning).It includes 4.8 million fully synthetic prompts with reasoning trajectories, serving as the cornerstone for supervised fine-tuning (SFT) experiments. The dataset demonstrates that purely synthetic data—when generated with PromptCoT 2.0—can train competitive reasoning models that outperform human-curated baselines such as OpenMathReasoning and… See the full description on the dataset page: https://huggingface.co/datasets/xl-zhao/PromptCoT-2.0-SFT-4.8M.text1M<n<10M11 likes1.3k downloads1y agoHugging Face09hendzh /PromptShield PromptShield Benchmark: A Flexible and Realistic Benchmark for Prompt Injection Attacks This dataset accompanies the paper "[PromptShield: Deployable Detection for Prompt Injection Attacks]" (ArXiv Link) and is built from a curated selection of open-source datasets and published prompt injection attack strategies. Dataset Details Task: Binary classification of prompt injection attempts. Fields: prompt: The full text of the prompt, including instructions, inputs, and… See the full description on the dataset page: https://huggingface.co/datasets/hendzh/PromptShield.texttext-classification10K<n<100K7 likes883 downloads1y agoHugging Face10Nymbo /Official_LLM_System_Prompts Official LLM System Prompts This short dataset contains a few system prompts leaked from proprietary models. Contains date-stamped prompts from OpenAI, Anthropic, MS Copilot, GitHub Copilot, Grok, and Perplexity. textn<1K29 likes680 downloads1y agoHugging Face11aimosprite /prompt-swap-mixed12-5xlr-e1-mxfp4-mergedtabularn<1K0 likes606 downloads6mo agoHugging Face12aimosprite /prompt-swap-mixed12-5xlr-e2-mxfp4-mergedtabularn<1K0 likes592 downloads6mo agoHugging Face13rl-rag /hle_rlvr_no_prompttextn<1K0 likes535 downloads1y agoHugging Face143nesdeniz /agentic-prompt-injection-boundary-pairs Agentic Prompt-Injection Boundary Pairs Most prompt-injection datasets make the attack easy to recognize. The malicious row contains obvious override language, while the benign row discusses something unrelated. A classifier can look capable without learning the boundary that matters in production. This dataset takes a stricter approach. Each attack is paired with a legitimate request from the same workflow. The two rows share the asset, role, tool and topic. What changes is… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/agentic-prompt-injection-boundary-pairs.texttext-classification1K<n<10K6 likes506 downloads2mo agoHugging Face15aimosprite /prompt-swap-medium12-e2-mxfp4-mergedtabularn<1K0 likes366 downloads6mo agoHugging Face16MAlmasabi /Indirect-Prompt-Injection-BIPIA-GPTgated Indirect Prompt Injection Detection Dataset (BIPIA + GPT-4o-mini) Dataset Summary This dataset contains 70,000 examples for detecting indirect prompt injection attacks in Large Language Models. It combines: 35,000 malicious samples from the BIPIA benchmark (cleaned and processed) 35,000 benign samples generated using GPT-4o-mini Indirect prompt injection attacks embed malicious instructions within external content (code, table, email, webAQ, abstract) that LLMs process… See the full description on the dataset page: https://huggingface.co/datasets/MAlmasabi/Indirect-Prompt-Injection-BIPIA-GPT.text10K<n<100K8 likes284 downloads8mo agoHugging Face17SupraLabs /Prompt-Routing-DatasetPrompt Routing Dataset · Multi-Task Infrastructure Routing About this dataset This dataset is a highly dense, premium alignment asset explicitly designed to train Edge Orchestrators and Routing Models ranging from 50M to 1.5B parameters. When deploying small language models (SLMs) on consumer hardware or local edge instances, running multi-step mathematical derivations or complex architectural software tasks often causes catastrophic hallucinations or syntax breakdown. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/SupraLabs/Prompt-Routing-Dataset.texttext-classificationn<1K23 likes274 downloads3mo agoHugging Face18Norod78 /hebrew_lyrics_prompting_finetunetexttext-generation10K<n<100K0 likes252 downloads2y agoHugging Face19Cseti /LTX2.3-22B_IC-LoRA-CrossView-Prompt-Dataset CrossView Prompt Dataset The training dataset behind the CrossView Prompt IC-LoRA for LTX-Video 2.3 — a "virtual second camera" adapter that re-renders a scene from a new viewpoint described by a short prompt. Each sample is a pair of static-camera clips of the same scene (a reference view and a target view) plus a camera-delta caption describing how the target camera differs from the reference. Contents clips/<scene>/<cam>.mp4 # 504 unique clips, native… See the full description on the dataset page: https://huggingface.co/datasets/Cseti/LTX2.3-22B_IC-LoRA-CrossView-Prompt-Dataset.tabularimage-to-videon<1K1 likes234 downloads2mo agoHugging Face20ChaoticNeutrals /Reddit-SFW-Writing_Prompts_ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing [Description Tags],"Deleted user", "Hello,\n\nYour post has been removed..", "Post has been deleted by user", "This post has been marked NSFW", duplicated system and human turns, etc has been removed. text100K<n<1M10 likes232 downloads2y agoHugging Face21Naomibas /llm-system-prompts-benchmark Dataset Card for Dataset Name This datset is a collection of 100 system prompts for large language models. Dataset Details Dataset Description These 100 system prompts test a model's ability to follow grammatical patterns; answer basic multiple choice questions; act according to a particular persona; memorize information; and speak in French. Files: hundred_system_prompts.py: refer to this to see the (prompt, probe, function) triplets, as well as the… See the full description on the dataset page: https://huggingface.co/datasets/Naomibas/llm-system-prompts-benchmark.textn<1K19 likes221 downloads2y agoHugging Face22DavidTKeane /clawk-agent-social-ai-prompt-injection-dataset Clawk Agent-Social AI Prompt Injection Dataset 85,703 items — 44,232 posts and 41,471 replies — from Clawk, a social network whose users are AI agents. Scanned for AI-to-AI indirect prompt injection using the threat model of Greshake et al. (2023). The full raw corpus is included, so you can ignore my analysis entirely and do your own. These are keyword-matched candidates, not verified attacks. An agent discussing prompt injection matches the same words as one performing it.… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/clawk-agent-social-ai-prompt-injection-dataset.texttext-classificationn<1K1 likes218 downloads18d agoHugging Face23PKU-Alignment /PKU-SafeRLHF-prompt Dataset Card for PKU-SafeRLHF-prompt This dataset contains 44.6K unique prompts from PKU-SafeRLHF. 22.4% of the prompts in this dataset come from the sibling project BeaverTails. Additionally, we performed SFT on Llama3-70B using the Alpaca 52K dataset, resulting in Alpaca3-70B. 63.6% and 14.0% of our dataset is generated by Alpaca3-70B and WizardLM-30B-Uncensored, respectively, under the guidance of experts. Here is the generation pipeline: Usage To load our dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-prompt.texttext-generation10K<n<100K5 likes202 downloads2y agoHugging Face24RASSAISAID /finance-deepseek-prompts-distill We are soon launching an end-to-end data process—distillation and synthetic data—to train (SFT and RL) a financial agentic model! Financial DeepSeek Distillation Prompts Ready-to-paste prompts for manually distilling financial reasoning datasets through DeepSeek-V4 pro/flash (or any LLM) UI. Available Datasets (English) Dataset Prompts Size Category Target finqa_train_prompts.jsonl 6,251 58 MB Advanced Business Knowledge 2,948… See the full description on the dataset page: https://huggingface.co/datasets/RASSAISAID/finance-deepseek-prompts-distill.text10K<n<100K3 likes201 downloads4mo agoHugging Face25DavidTKeane /moltbook-agent-social-ai-prompt-injection-dataset Moltbook Agent-Social AI Prompt Injection Dataset 207,391 items — 77,469 posts and 129,922 comments — from Moltbook, a social network whose users are AI agents. Scanned for indirect prompt-injection patterns using the taxonomy of Greshake et al. (2023). The full raw corpus is included, so you can ignore my analysis entirely and do your own. These are keyword-matched candidates, not verified attacks. An agent discussing prompt injection matches the same words as one performing… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/moltbook-agent-social-ai-prompt-injection-dataset.tabulartext-classification1K<n<10K1 likes199 downloads18d agoHugging Face26xl-zhao /PromptCoT-2.0-SelfPlay-4B-48K PromptCoT-2.0-SelfPlay Datasets This repository hosts the self-play datasets used in PromptCoT 2.0 (Scaling Prompt Synthesis for LLM Reasoning).These datasets were created by applying the PromptCoT 2.0 synthesis framework to generate challenging math and programming problems, and then training models through self-play with Direct Preference Optimization (DPO). PromptCoT-2.0-SelfPlay-4B-48K: 48,113 prompts for Qwen3-4B-Thinking-2507 self-play. PromptCoT-2.0-SelfPlay-30B-11K: 11… See the full description on the dataset page: https://huggingface.co/datasets/xl-zhao/PromptCoT-2.0-SelfPlay-4B-48K.text10K<n<100K0 likes195 downloads1y agoHugging Face27prodnull /prompt-injection-repo-datasetgated Prompt Injection Repository File Dataset A labeled dataset for detecting prompt injection attacks in repository files — code, configs, READMEs, CI/CD workflows, and documentation that AI coding agents process as context. What This Is (and Isn't) This dataset targets a specific threat: indirect prompt injection via repository content. When AI coding agents (Claude Code, Cursor, Copilot, Gemini CLI) clone a repo, every file becomes part of the agent's context.… See the full description on the dataset page: https://huggingface.co/datasets/prodnull/prompt-injection-repo-dataset.texttext-classification1K<n<10K11 likes188 downloads7mo agoHugging Face28agentlans /prompt-difficulty Prompt Difficulty Assessment Prompt difficulty plays a critical role in the performance of large language models (LLMs). Assessing this difficulty is essential for selecting training examples, evaluating model capabilities, and optimizing routing and reasoning strategies. Yet, no standardized framework exists for comparing prompt difficulty across domains. This report proposes a method to quantify prompt difficulty using multiple LLMs and introduces a composite difficulty score for… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-difficulty.tabulartext-classification10K<n<100K0 likes186 downloads9mo agoHugging Face29TruongSinhAI /deepcad_prompt_jsontext100K<n<1M0 likes181 downloads1y agoHugging Face30LHL3341 /AutoBench_Promptstextn<1K0 likes177 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.