datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tm-system_promptdrh-System-Prompt-processedOfficial_LLM_System_Prompts
Official LLM System Prompts
This short dataset contains a few system prompts leaked from proprietary models. Contains date-stamped prompts from OpenAI, Anthropic, MS Copilot, GitHub Copilot, Grok, and Perplexity.
OpenThoughts3-456k-no-cot-with-olmo-system-promptsystem-prompt-leakage
System Prompt Leakage Dataset
Overview
The System Prompt Leakage Dataset offers a collection of synthetic prompts and model responses, specifically designed to help detect and manage instances of system prompt leakage. In modern applications of large language models (LLMs), safeguarding sensitive or proprietary system instructions from being exposed in responses is critical. This dataset provides a diverse set of real-world-inspired examples for developing and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/gabrielchua/system-prompt-leakage.llm-system-prompts-benchmark
Dataset Card for Dataset Name
This datset is a collection of 100 system prompts for large language models.
Dataset Details
Dataset Description
These 100 system prompts test a model's ability to follow grammatical patterns; answer basic multiple choice questions; act according to a particular persona; memorize information; and speak in French.
Files:
hundred_system_prompts.py: refer to this to see the (prompt, probe, function) triplets, as well as the… See the full description on the dataset page: https://huggingface.co/datasets/Naomibas/llm-system-prompts-benchmark.system-prompt-sft-50k
System Prompt Diversity SFT (50K)
50,000 conversations in ShareGPT format where the assistant correctly follows diverse system prompt personas and constraints.
Motivation
A model that ignores system prompts is useless in production. The most common alignment failure in deployed LLMs is drift from system-level instructions: breaking persona, discussing off-topic subjects, ignoring tone or format constraints, and failing role-specific guardrails. This dataset trains… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/system-prompt-sft-50k.rl_rag_sqa_searcharena_rubrics_web_augmented_outcome_with_new_mcp_system_prompttulu-3-sft-mixture-olmo-2-system-promptrl_rag_sqa_searcharena_rubrics_web_augmented_rubrics_only_with_new_mcp_system_promptconfigurable-system-prompt-multitask
Configurable System Prompt Multi-task Dataset 🛞
We release the synthetic dataset for the multi-task experiments from the paper "Configurable Safety Tuning of Language Models with Synthetic Preference Data", https://huggingface.co/papers/2404.00495. This dataset has two sources for the examples:
Self-critique on a safety task from Harmful Behaviours, using the SOLAR-Instruct model. It employs two system prompts to learn the different behaviors:
You are a helpful yet harmless… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/configurable-system-prompt-multitask.daily_dialog_with_system_promptssystem-prompt-reasoning-traces
System-Prompt Reasoning Traces
A novel dataset combining system prompt adherence with structured internal reasoning traces, built on findings from 14+ research papers.
🔬 Research Foundation
This dataset is the first to systematically combine system prompt diversity with structured reasoning traces. It incorporates findings from:
Paper
Key Finding
How We Use It
Sky-T1 (Berkeley, 2025)
Structure > content in reasoning traces — wrong answers with good structure… See the full description on the dataset page: https://huggingface.co/datasets/Michael-Kozu/system-prompt-reasoning-traces.HuggingChat-AI-Assistants-Deleted-System-Promptstulu-3-sft-personas-math-grade-o3-with-olmo-system-promptsystem-prompts-autonomous-agentswildchat_perturbed_6000_replaced_no_keyword-with-olmo-system-promptallenai_tulu-3-sft-olmo-2-mixture-0225__system_prompt_2023various_RP_system_promptsCollection of various system prompts for RP. Feel free to contribute more by opening a discussion.
ChuckMcSneed-interesting:
my currently favorite system prompt
includes Orwells writing rules
+writes in non-boring style
+more realistic reactions
+eliminates a lot of GPTslop
-writing style is not for everyone
-complains more, but still does what is requested
-sometimes reddit-like
ChuckMcSneed-multistyle
List of various styles
All of them tested with examples
Ranges from good to shit… See the full description on the dataset page: https://huggingface.co/datasets/ChuckMcSneed/various_RP_system_prompts.Gemini-2_SystemPromptmagpie-llama-3.1-405b-instruct-fp8-with-system-prompt-per-category
Dataset Card for magpie-llama-3.1-405b-instruct-fp8-with-system-prompt-per-category
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/gabrielmbmb/magpie-llama-3.1-405b-instruct-fp8-with-system-prompt-per-category/raw/main/pipeline.yaml"
or… See the full description on the dataset page: https://huggingface.co/datasets/gabrielmbmb/magpie-llama-3.1-405b-instruct-fp8-with-system-prompt-per-category.hermes3_special_system_prompts
Hermes 3 SFT Special System Prompts
Version of the NousResearch/Hermes-3-Dataset dataset used in AMALIA's Supervised Fine-Tuning stage.
This subset was manually selected in order to retain only entries containing custom system prompts that substantially modify the model’s behavior. It were also built splits with the highest quality entries and a translated part of those entries to European Portuguese. Both the quality classification and translation were done using… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/hermes3_special_system_prompts.System-Prompt-Instruction-Real-world-Implementation-Training-set
SPIRIT Dataset (System Prompt Instruction Real-world Implementation Training-set)
Dataset Summary
SPIRIT is a high-quality system prompt instruction dataset designed to enhance language models' ability to follow complex system prompts. The dataset comprises real-world system prompts collected from GitHub repositories and synthetically generated conversations, specifically curated to improve system prompt adherence in large language models.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/EricLu/System-Prompt-Instruction-Real-world-Implementation-Training-set.hermes3_system_promptsno_robots_converted-with-olmo-system-prompttulu-3-sft-coconot-regenerated-with-olmo-system-prompttulu_v3.9_wildjailbreak_decontaminated_50k-with-olmo-system-promptvicgalle_configurable-system-prompt-multitask-PreferenceShareGPTorca_dpo_pairs_no_system_prompttulu-3-sft-personas-math-o3-with-olmo-system-prompt
