CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /real-toxicity-prompts Dataset Card for Real Toxicity Prompts Dataset Summary RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models. Languages English Dataset Structure Data Instances Each instance represents a prompt and its metadata: { "filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt", "begin":340, "end":564, "challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/allenai/real-toxicity-prompts.tabular10K<n<100K123 likes20k downloads4y agoHugging Face02Gryphe /ChatGPT-4o-Writing-Prompts ChatGPT-4o Writing Prompts This is a dataset containing 3746 short stories, generated with OpenAI's chatgpt-4o-latest model and using Reddit's Writing Prompts subreddit as a source. Each sample is generally between 6000-8000 characters long. These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres. Note that I did not touch the Markdown ChatGPT-4o produced by itself to enrich its output, as I very much enjoy the added flavour… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts.texttext-generation1K<n<10K36 likes7.1k downloads2y agoHugging Face03artificialguybr /veo3-video-prompts Veo 3 Video Generation Dataset English | Português do Brasil English Summary A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant. Videos: 5,811 Input images: 1,354 Configurations: 6 Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.imagetext-to-video1K<n<10K0 likes5.3k downloads1mo agoHugging Face04hendzh /PromptShield PromptShield Benchmark: A Flexible and Realistic Benchmark for Prompt Injection Attacks This dataset accompanies the paper "[PromptShield: Deployable Detection for Prompt Injection Attacks]" (ArXiv Link) and is built from a curated selection of open-source datasets and published prompt injection attack strategies. Dataset Details Task: Binary classification of prompt injection attempts. Fields: prompt: The full text of the prompt, including instructions, inputs, and… See the full description on the dataset page: https://huggingface.co/datasets/hendzh/PromptShield.texttext-classification10K<n<100K7 likes873 downloads1y agoHugging Face05Nymbo /Official_LLM_System_Prompts Official LLM System Prompts This short dataset contains a few system prompts leaked from proprietary models. Contains date-stamped prompts from OpenAI, Anthropic, MS Copilot, GitHub Copilot, Grok, and Perplexity. textn<1K29 likes736 downloads1y agoHugging Face06aimosprite /prompt-swap-mixed12-5xlr-e1-mxfp4-mergedtabularn<1K0 likes588 downloads6mo agoHugging Face07aimosprite /prompt-swap-mixed12-5xlr-e2-mxfp4-mergedtabularn<1K0 likes579 downloads6mo agoHugging Face08aimosprite /prompt-swap-medium12-e2-mxfp4-mergedtabularn<1K0 likes382 downloads6mo agoHugging Face09ChaoticNeutrals /Reddit-SFW-Writing_Prompts_ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing [Description Tags],"Deleted user", "Hello,\n\nYour post has been removed..", "Post has been deleted by user", "This post has been marked NSFW", duplicated system and human turns, etc has been removed. text100K<n<1M10 likes251 downloads2y agoHugging Face10Naomibas /llm-system-prompts-benchmark Dataset Card for Dataset Name This datset is a collection of 100 system prompts for large language models. Dataset Details Dataset Description These 100 system prompts test a model's ability to follow grammatical patterns; answer basic multiple choice questions; act according to a particular persona; memorize information; and speak in French. Files: hundred_system_prompts.py: refer to this to see the (prompt, probe, function) triplets, as well as the… See the full description on the dataset page: https://huggingface.co/datasets/Naomibas/llm-system-prompts-benchmark.textn<1K19 likes222 downloads2y agoHugging Face11jayseanbrambila /minimax-h3-video-prompts MiniMax H3 Video Prompts A small, curated collection of 50 structured prompts for text-to-video and image-to-video workflows. It covers cinematic scenes, characters, animation, nature, architecture, product shots, food, social video, and fantasy environments. Use the prompts to create video Copy a prompt from the dataset, adapt it to your idea, then generate the finished video online. Create an AI video with MiniMax3.org → Dataset details… See the full description on the dataset page: https://huggingface.co/datasets/jayseanbrambila/minimax-h3-video-prompts.texttext-to-videon<1K0 likes192 downloads25d agoHugging Face12LHL3341 /AutoBench_Promptstextn<1K0 likes184 downloads9mo agoHugging Face13RASSAISAID /finance-deepseek-prompts-distill We are soon launching an end-to-end data process—distillation and synthetic data—to train (SFT and RL) a financial agentic model! Financial DeepSeek Distillation Prompts Ready-to-paste prompts for manually distilling financial reasoning datasets through DeepSeek-V4 pro/flash (or any LLM) UI. Available Datasets (English) Dataset Prompts Size Category Target finqa_train_prompts.jsonl 6,251 58 MB Advanced Business Knowledge 2,948… See the full description on the dataset page: https://huggingface.co/datasets/RASSAISAID/finance-deepseek-prompts-distill.text10K<n<100K3 likes165 downloads4mo agoHugging Face14lautaschiaffino /wikiprompt-prompts Wikiprompt Prompts The full prompt catalog of wikiprompt.org, the free prompt encyclopedia: 115,000+ curated AI prompts with structured metadata (model, media type, style, quality assessment, keywords) and links to the live pages. Configs default - the FULL catalog (wikiprompt_prompts.jsonl), refreshed monthly. This is the row count you should compare against wikiprompt.org. year-YYYY / YYYY-MM - stable snapshots per closed year and per closed month of the… See the full description on the dataset page: https://huggingface.co/datasets/lautaschiaffino/wikiprompt-prompts.text100K<n<1M1 likes162 downloads6d agoHugging Face15rishi-1001 /webcode2m-natural-promptstext1K<n<10K1 likes138 downloads1y agoHugging Face16selimaktas /turkish-flow-drafter-prompts Turkish prompts for Chained-Flow drafter training Chat-templated Turkish prompts used to train and evaluate the Turkish Flow-Drafter checkpoints for Qwen/Qwen3.5-4B / 9B / 27B. Prompts only — no completions. A drafter is trained on the target model's own hidden states, so continuations are generated locally by running the target over these prompts. Nothing here is a model output. split rows what it is v1/ 29,100 train + 300 holdout the mixture the released Turkish… See the full description on the dataset page: https://huggingface.co/datasets/selimaktas/turkish-flow-drafter-prompts.texttext-generation10K<n<100K0 likes130 downloads20d agoHugging Face17k-mktr /llm_eval_promptstextquestion-answering1K<n<10K1 likes129 downloads2y agoHugging Face18guicybercode /japan-math-philosophy-prompts Japan Math Philosophy Prompts Microdataset autoral com problemas que combinam matemática e reflexão filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em pt-BR, en e ja e mantida integralmente no split train. Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.textquestion-answeringn<1K0 likes119 downloads27d agoHugging Face19cfahlgren1 /llama-3.1-awesome-chatgpt-prompts Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/llama-3.1-awesome-chatgpt-prompts.tabularn<1K7 likes116 downloads2y agoHugging Face20agentlans /allenai-WildChat-4.8M-prompts allenai/WildChat-4.8M English Prompts Dataset Summary This dataset contains real user-submitted prompts to ChatGPT, extracted from the English portion of the allenai/WildChat-4.8M collection. It serves as a large-scale resource for analyzing user intent, conversational diversity, and prompt engineering patterns. Files en_prompts: All English-language first messages from user conversations. Each record represents the first user prompt. Exact duplicates are… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/allenai-WildChat-4.8M-prompts.texttext-generation1M<n<10M0 likes113 downloads11mo agoHugging Face21lamm-mit /gemma4-materials-mechanism-prompts Gemma 4 Materials-Mechanism Prompt Corpus This dataset collects the exact scientific prompts and registered prompt metadata used in “Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model” by Markus J. Buehler. It is organized as 21 Hugging Face configurations so that historical development prompts, frozen evaluations, falsification tests, and exploratory follow-ups are not pooled into one ambiguous table. The release is a prompt and… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-materials-mechanism-prompts.textquestion-answering1K<n<10K0 likes113 downloads2mo agoHugging Face22ChaoticNeutrals /Reddit-NSFW-Writing_Prompts_ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing [Description Tags],"Deleted user", "Hello,\n\nYour post has been removed..", "Post has been deleted by user", "This post has been marked NSFW", duplicated system and human turns, etc has been removed. text1K<n<10K17 likes109 downloads2y agoHugging Face23Delta-Vector /Tauri-Opus-Accepted-GPT-Rejected-Opus-Writing-Promptstext1K<n<10K3 likes107 downloads1y agoHugging Face24agentlans /prompt-safety-scores Composite Safety Scoring for Prompts Using Multiple LLM Annotations Introduction Evaluating the safety of prompts is essential but challenging. Existing approaches often depend on predefined categories, which can be circumvented by new jailbreaks or attacks. Additionally, different tasks may require different safety thresholds. This study explores using large language models (LLMs) themselves to annotate prompt safety. By combining these annotations, a continuous safety… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-safety-scores.tabulartext-classification10K<n<100K0 likes105 downloads6mo agoHugging Face25allenai /tulu-2.5-prompts Tulu 2.5 Prompts Dataset This dataset contains the set of prompts used to train the PPO models described in Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback. This contains only the prompts used during the PPO training. Dataset Details The description of each prompt goes as follows: gsm8k_prompts: Prompts taken from the GSM8k train split. ultrafeedback_prompts: The prompts from the cleaned UltraFeedback dataset. math_prompts:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-2.5-prompts.text100K<n<1M4 likes104 downloads2y agoHugging Face26endjin /120-years-of-olympic-history-athletes-and-results-promptstext1M<n<10M0 likes103 downloads2y agoHugging Face27Perflow-Shuai /cfg-sensitive-video-prompts CFG-Sensitive Video Prompts A provenance-first, text-only prompt suite for comparing classifier-free guidance behavior in video generation. It contains 22 English prompts: existing_visualized: 6 prompts used by the local Wan2.2-TI2V-5B CFG-only LoRA scale-sweep report. cfg_sensitive_candidate: 16 additional prompts selected because color/attribute binding, numeracy, readable text, spatial relations, temporal change, fluid motion, weather, camera motion, multi-subject… See the full description on the dataset page: https://huggingface.co/datasets/Perflow-Shuai/cfg-sensitive-video-prompts.tabulartext-to-videon<1K0 likes102 downloads16d agoHugging Face28Houzeric /french-prompts-and-questions Dataset français Prompt / Answer (multi-sources) 📌 Description Ce dataset est un jeu de données en français au format prompt / answer, conçu pour l’entraînement, le fine-tuning et l’évaluation de modèles de langage (LLM). Il résulte de la fusion et de la normalisation de plusieurs datasets publics, afin de proposer un format simple, homogène et facilement exploitable. Chaque entrée contient : prompt : question ou instruction utilisateur answer : réponse attendue type :… See the full description on the dataset page: https://huggingface.co/datasets/Houzeric/french-prompts-and-questions.text100K<n<1M1 likes93 downloads9mo agoHugging Face29Roman1111111 /prompts-for-claude-opus-4.6text10K<n<100K6 likes93 downloads6mo agoHugging Face30RASSAISAID /finqa-deepseek-prompts FinQA → DeepSeek Distillation Prompts Ready-to-paste prompts for manually distilling FinQA through DeepSeek-R1 (or any LLM) UI. Files File Samples Size Description finqa_train_prompts.jsonl 6,251 58 MB Training split prompts finqa_dev_prompts.jsonl 883 8 MB Validation split prompts finqa_test_prompts.jsonl 1,147 11 MB Test split prompts finqa_to_prompts.py — 5 KB Converter script (raw FinQA JSON → prompts) view_prompts_for_ui.py — 3 KB Terminal… See the full description on the dataset page: https://huggingface.co/datasets/RASSAISAID/finqa-deepseek-prompts.text1K<n<10K1 likes92 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.