CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Gryphe /ChatGPT-4o-Writing-Prompts ChatGPT-4o Writing Prompts This is a dataset containing 3746 short stories, generated with OpenAI's chatgpt-4o-latest model and using Reddit's Writing Prompts subreddit as a source. Each sample is generally between 6000-8000 characters long. These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres. Note that I did not touch the Markdown ChatGPT-4o produced by itself to enrich its output, as I very much enjoy the added flavour… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts.texttext-generation1K<n<10K36 likes7.2k downloads2y agoHugging Face02guicybercode /japan-math-philosophy-prompts Japan Math Philosophy Prompts Microdataset autoral com problemas que combinam matemática e reflexão filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em pt-BR, en e ja e mantida integralmente no split train. Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.textquestion-answeringn<1K0 likes119 downloads28d agoHugging Face03agentlans /allenai-WildChat-4.8M-prompts allenai/WildChat-4.8M English Prompts Dataset Summary This dataset contains real user-submitted prompts to ChatGPT, extracted from the English portion of the allenai/WildChat-4.8M collection. It serves as a large-scale resource for analyzing user intent, conversational diversity, and prompt engineering patterns. Files en_prompts: All English-language first messages from user conversations. Each record represents the first user prompt. Exact duplicates are… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/allenai-WildChat-4.8M-prompts.texttext-generation1M<n<10M0 likes106 downloads11mo agoHugging Face04selimaktas /turkish-flow-drafter-prompts Turkish prompts for Chained-Flow drafter training Chat-templated Turkish prompts used to train and evaluate the Turkish Flow-Drafter checkpoints for Qwen/Qwen3.5-4B / 9B / 27B. Prompts only — no completions. A drafter is trained on the target model's own hidden states, so continuations are generated locally by running the target over these prompts. Nothing here is a model output. split rows what it is v1/ 29,100 train + 300 holdout the mixture the released Turkish… See the full description on the dataset page: https://huggingface.co/datasets/selimaktas/turkish-flow-drafter-prompts.texttext-generation10K<n<100K0 likes103 downloads21d agoHugging Face05ytu-ce-cosmos /turkish-flow-drafter-prompts GitHub repo · Technical blog · Model collection Turkish prompts for Chained-Flow drafter training Chat-templated Turkish prompts used to train and evaluate the Turkish Flow-Drafter checkpoints for Qwen/Qwen3.5-4B / 9B / 27B. Prompts only — no completions. A drafter is trained on the target model's own hidden states, so continuations are generated locally by running the target over these prompts. Nothing here is a model output. split rows what it is v1/… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/turkish-flow-drafter-prompts.texttext-generation10K<n<100K0 likes95 downloads9d agoHugging Face06selimaktas /english-flow-drafter-prompts English prompts for Chained-Flow drafter training Chat-templated English prompts used to train the English Flow-Drafter checkpoints for Qwen/Qwen3.5-4B / 9B / 27B. Prompts only — no completions. A drafter is trained on the target model's own hidden states, so continuations are generated locally by running the target over these prompts. Nothing here is a model output. split rows prompt tokens what it is v1/ 19,672 train + 500 holdout 1,518,620 the original mixture… See the full description on the dataset page: https://huggingface.co/datasets/selimaktas/english-flow-drafter-prompts.texttext-generation10K<n<100K0 likes83 downloads15d agoHugging Face07ai-mitra /prompt-slimmer-slm Prompt Slimmer SLM — Demo Dataset Synthetic examples for experimenting with prompt rewriting and sentence selection. Exported without changing the examples or their original splits from the shared GitHub codebase. Model · Project page Configuration Train Validation Test Purpose rewrites-expanded (default) 41 2 2 Expanded rewriting dataset: 45 examples rewrites 9 2 2 Original dataset used by the first adapter selector 256 64 64 KEEP/DROP labels for source spans… See the full description on the dataset page: https://huggingface.co/datasets/ai-mitra/prompt-slimmer-slm.texttext-generationn<1K0 likes81 downloads9d agoHugging Face08ytu-ce-cosmos /english-flow-drafter-prompts GitHub repo · Technical blog · Model collection English prompts for Chained-Flow drafter training Chat-templated English prompts used to train the English Flow-Drafter checkpoints for Qwen/Qwen3.5-4B / 9B / 27B. Prompts only — no completions. A drafter is trained on the target model's own hidden states, so continuations are generated locally by running the target over these prompts. Nothing here is a model output. split rows prompt tokens what it is v1/… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/english-flow-drafter-prompts.texttext-generation10K<n<100K0 likes61 downloads9d agoHugging Face09guicybercode /iceland-tech-christian-ethics-prompts Fictional Icelandic Landscapes, Technology and Christian Ethics Prompts This microdataset contains 24 original discussion prompts arranged as 12 parallel pt-BR/English pairs. Each explicitly fictional scenario combines a landscape motif inspired by Iceland, a technology-governance dilemma, and concepts that may be explored through Christian ethics. The records do not describe real Icelandic institutions, policies, communities, or practices, and they do not claim that Christians… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/iceland-tech-christian-ethics-prompts.texttext-generationn<1K0 likes58 downloads28d agoHugging Face10NikoThePig /internet-prompts-benchmark Internet Prompts Benchmark Viral internet prompts, memes, and tests that AI historically failed at. Popular ones like counting letters in the word strawberry and nicher ones that test other important capabilities.Mostly made by GPT-5.6 Sol. It is designed for many types of models to participate, small and large, not only transformers. It has prompts from the early days of AI to the very latest. It will be actively updated to preserve various prompts for as long as I can afford… See the full description on the dataset page: https://huggingface.co/datasets/NikoThePig/internet-prompts-benchmark.texttext-generationn<1K0 likes54 downloads15d agoHugging Face11Nbardy /diverse-svg-prompts Diverse SVG Prompts Diverse SVG Prompts is a public collection of 20,000 high-quality, generated and filtered English briefs for SVG and vector-graphics generation. It contains 18,000 general illustration prompts and 2,000 lettering prompts. Schema The dataset intentionally has only two columns: prompt: the complete visual brief. type_tags: a list of category, author-model, and processing tags. Example: { "prompt": "A moonlit mechanical heron..."… See the full description on the dataset page: https://huggingface.co/datasets/Nbardy/diverse-svg-prompts.texttext-generation10K<n<100K0 likes52 downloads28d agoHugging Face12gray311 /PromptSD PromptSD Training and evaluation data for PromptSD, an on-policy soft-prompt-teacher distillation method. The release covers the four target tasks used in the paper. Every example carries a <reasoning>...</reasoning> chain followed by a <answer>...</answer> span, so the data can be used directly for reasoning-supervised SFT, distillation, or RLVR. Configurations Config (config_name) Task Source / format Train Validation Test science Science MCQ 4-way… See the full description on the dataset page: https://huggingface.co/datasets/gray311/PromptSD.textquestion-answering1K<n<10K0 likes51 downloads4mo agoHugging Face13LaelaZorana /synthetic-instruction-promptsgated Synthetic Instruction Prompts (8 domains) Most synthetic prompt sets are a black box. You get a pile of prompts and no idea whether they're actually varied or just the same three sentences wearing different nouns. This one is graded, and the grade is on the card. It's 2,829 instruction-style prompts across eight domains, generated with SynthKit and then scored by the same tool. The prompts carry no answers. Think of them as seed prompts: you feed them to a model to bootstrap… See the full description on the dataset page: https://huggingface.co/datasets/LaelaZorana/synthetic-instruction-prompts.texttext-generation1K<n<10K0 likes51 downloads10d agoHugging Face14kishormorol /promptlean-prompts PromptLean Prompts 120 prompts across 14 categories, each in three variants: Lean, Balanced, and Max Quality. Averaged across the library, the Lean variant uses about 86% fewer tokens than Max Quality. This is the data behind PromptLean. The idea Most published prompts are overengineered. A code review does not need 200 tokens of preamble, but some tasks genuinely do earn the extra context. Keeping all three variants side by side makes that tradeoff explicit and… See the full description on the dataset page: https://huggingface.co/datasets/kishormorol/promptlean-prompts.texttext-generationn<1K0 likes49 downloads8d agoHugging Face15vinci00 /ministral-3-benchmark-prompts Ministral 3 MLX benchmark prompts This tiny dataset contains the four fixed prompts used by the reproducible smoke benchmark for the Ministral 3 MLX 4-bit model. It is a benchmark fixture, not a training or fine-tuning dataset. Schema Each JSONL row contains: id: stable case identifier; language: prompt language; prompt: exact input sent to the model; expected_keywords: lowercase substrings used by the smoke check. The benchmark uses greedy decoding and checks… See the full description on the dataset page: https://huggingface.co/datasets/vinci00/ministral-3-benchmark-prompts.texttext-generationn<1K0 likes46 downloads6d agoHugging Face16nmsofficial /Manim-8600-Prompts ManimCoder Prompts An English prompt dataset for generating Manim Community scenes and related engineering tasks. Dataset size Source collection: 9,000 records Deduplicated release: 8,600 prompts Removed structural duplicates: 400 The removed records came from one source section where each of 100 primary 3D objectives had been repeated five times with only the camera or reveal instruction changed. One variant per primary objective was retained.… See the full description on the dataset page: https://huggingface.co/datasets/nmsofficial/Manim-8600-Prompts.texttext-generation1K<n<10K0 likes43 downloads2mo agoHugging Face17vislupus /alpaca-bulgarian-jokes-multilingual-prompts Bulgarian Jokes Dataset Overview The Bulgarian Jokes Dataset is a collection of Bulgarian-language jokes gathered and prepared for use in training and fine-tuning natural language processing (NLP) models. This dataset is designed to help researchers and developers build models capable of understanding and generating humorous content in Bulgarian. Dataset Structure The dataset is structured in a format suitable for NLP training and fine-tuning tasks, such as the… See the full description on the dataset page: https://huggingface.co/datasets/vislupus/alpaca-bulgarian-jokes-multilingual-prompts.texttext-generation10K<n<100K1 likes39 downloads2y agoHugging Face18rupeshreddypapa /nepi-prompts-dataset NEPI: Narrative-Embedded Prompt Injection Dataset (Sanitized) Dataset Summary This dataset contains 4,000 sanitized prompts designed for research on prompt injection vulnerabilities in Large Language Models (LLMs).It introduces and supports evaluation of a novel attack class called Narrative-Embedded Prompt Injection (NEPI), where adversarial intent is embedded inside coherent fictional narratives, dialogues, or persona-driven roleplay prompts. Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/rupeshreddypapa/nepi-prompts-dataset.texttext-generation1K<n<10K0 likes38 downloads7d agoHugging Face19harithoppil /ml-swe-prompts ML SWE Prompts Unified collection of ML/training-related software engineering prompts for OPD distillation training. All prompts are in English. Filtered to core ML repos: huggingface (1,058), numpy (937), Lightning-AI (377), ray-project (342). Excludes pandas-dev, qiskit, open-mmlab, scipy, tensorflow, spaCy. Splits Config Source Rows Description all Combined 6,220 All prompts combined swe_bench_ml SWE-bench train 2,714 Problem statements from core ML repos… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/ml-swe-prompts.texttext-generation1K<n<10K0 likes35 downloads5mo agoHugging Face20kth8 /system_prompts_SuperGPQA-26000xSFT system prompts dataset generated using openai/gpt-oss-120b and m-a-p/SuperGPQA dataset. Each instance follows this format: { "uuid": "000192f411a04f13858d69834a44ae01", "messages": [ { "role": "system", "content": "You are a system prompt generator."}, { "role": "user", "content": "Write a system prompt that defines an AI researcher who is a leading authority in Science, specifically in Physics and Quantum Mechanics." }, { "role":… See the full description on the dataset page: https://huggingface.co/datasets/kth8/system_prompts_SuperGPQA-26000x.texttext-generation10K<n<100K0 likes34 downloads6mo agoHugging Face21YILMAZB1 /vidiary-reflective-prompts ViDiary Reflective Prompts & Emotional Taxonomy Dataset This open dataset contains foundational reflective journaling prompts and emotional sentiment taxonomy used in the development of ViDiary — the AI-powered voice and video journal with dual-PIN Decoy Vault. 🎙️ About ViDiary ViDiary is an innovative mobile application engineered to solve the #1 psychological hurdle in personal self-care: Bedtime Typing Fatigue. Research shows that over 80% of… See the full description on the dataset page: https://huggingface.co/datasets/YILMAZB1/vidiary-reflective-prompts.texttext-generationn<1K0 likes34 downloads16d agoHugging Face22kth8 /system_prompts_Jobs-20000xSFT system prompts dataset generated using openai/gpt-oss-120b and Faker jobs library. Each instance follows this format: { "uuid": "7ef8e7a637934d1d9ddf0856ba6bda98", "messages": [ { "role": "system", "content": "You are a system prompt generator." }, { "role": "user", "content": "Design a system prompt for an AI assistant that excels at answering advanced questions about Outdoor activities/education manager." }, { "role": "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/kth8/system_prompts_Jobs-20000x.texttext-generation10K<n<100K0 likes33 downloads6mo agoHugging Face239mark9 /llm-redteam-owasp-prompts LLM Red-Team Prompts — OWASP LLM Top 10 A curated dataset of 150 adversarial red-team prompts for evaluating the safety and robustness of large language models, mapped to the OWASP LLM Top 10. Every prompt is a real payload extracted directly from the open-source llm-safety-auditor project — none are fabricated. The dataset combines two sources from that project: 50 hand-curated attack templates (attack_library) — 10 per attack category. 100 mutation-engine variants… See the full description on the dataset page: https://huggingface.co/datasets/9mark9/llm-redteam-owasp-prompts.texttext-classificationn<1K1 likes33 downloads3mo agoHugging Face24slavazeph /xio-compliance-brain-triad-prompts XIO Compliance Brain — Triad Reviewer Prompts Reusable system prompts for running a multi-voice compliance debate against the same matter — the heart of XIO Compliance Brain's "Triad Review Engine" pattern. This dataset extracts the production prompts from the open-source compliance-AI hackathon branch so others can replicate the Triad pattern (three reviewer voices + synthesis + optional Round 2) on any LLM that follows OpenAI-compatible chat APIs. What's in this… See the full description on the dataset page: https://huggingface.co/datasets/slavazeph/xio-compliance-brain-triad-prompts.texttext-generationn<1K0 likes31 downloads5mo agoHugging Face25kondasviktor /vcl-ai-coding-prompts VCL AI Coding Power Prompts 50 battle-tested prompts for Claude Code, Codex, Gemini CLI, and Cursor — by Vibe Coder's Life. Free catalog for vibe coders. Replace {{PLACEHOLDERS}} with your facts. Not the paid Apify Playbook prompt pack (those stay private). Load from datasets import load_dataset ds = load_dataset("kondasviktor/vcl-ai-coding-prompts", "prompts") print(ds["train"][0]["title"]) Columns Column Description id Stable id… See the full description on the dataset page: https://huggingface.co/datasets/kondasviktor/vcl-ai-coding-prompts.texttext-generationn<1K0 likes28 downloads10d agoHugging Face26Vaibhav-GOAT /nepi-prompts-dataset NEPI: Narrative-Embedded Prompt Injection Dataset (Sanitized) Dataset Summary This dataset contains 4,000 sanitized prompts designed for research on prompt injection vulnerabilities in Large Language Models (LLMs).It introduces and supports evaluation of a novel attack class called Narrative-Embedded Prompt Injection (NEPI), where adversarial intent is embedded inside coherent fictional narratives, dialogues, or persona-driven roleplay prompts. Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/Vaibhav-GOAT/nepi-prompts-dataset.texttext-generation1K<n<10K0 likes23 downloads8mo agoHugging Face27procodec /sarvam-30b-audit-prompts Sarvam-30B Responsible-AI Audit — Pre-Registered Prompt Manifest 120 prompts across 5 categories, sampled deterministically (seed = 42) and pre-registered as the eval contract for a public responsible-AI audit of Sarvam-30B, India's sovereign-built reasoning LLM. This dataset is the eval contract committed to git before any prompt was sent to the model. Reviewers can verify every prompt by going to the cited source and pulling that exact row. Composition #… See the full description on the dataset page: https://huggingface.co/datasets/procodec/sarvam-30b-audit-prompts.texttext-generationn<1K0 likes22 downloads4mo agoHugging Face28tim-gabie /am-deepseek-r1-distilled-prompts-1.4m AM DeepSeek R1 Distilled Prompts 1.4M This dataset contains prompt-only rows extracted from a-m-team/AM-DeepSeek-R1-Distilled-1.4M. Extraction For each source JSONL row, every message with role == "user" was emitted as one prompt row. Assistant responses, reasoning traces, answers, and source metadata were not included. Source files: am_0.5M.jsonl.zst am_0.9M.jsonl.zst Extraction results: Input rows: 1,400,000 Output prompt rows: 1,400,000 JSON parse errors: 0… See the full description on the dataset page: https://huggingface.co/datasets/tim-gabie/am-deepseek-r1-distilled-prompts-1.4m.texttext-generation1M<n<10M0 likes21 downloads3mo agoHugging Face29TheFatBlue /llm-attacked-prompts-clm LLM-Paraphrased Adversarial Prompts LLM-paraphrased adversarial prompts for three code-generation benchmarks (MBPP+, HumanEval+, CanItEdit), used by RobustEval-CLM's LLMParaphraseAttack. Each row corresponds to one task in the source benchmark and carries the original prompt alongside an adversarial rewrite produced by an LLM under a BERTScore faithfulness constraint. Configs config source benchmark rewrite surface mbpp MBPP+ line 1 of the 4-line prompt… See the full description on the dataset page: https://huggingface.co/datasets/TheFatBlue/llm-attacked-prompts-clm.texttext-generationn<1K0 likes20 downloads5mo agoHugging Face30PJMixers-Dev /Gryphe_ChatGPT-4o-Writing-Prompts-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT Gryphe_ChatGPT-4o-Writing-Prompts-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT Gryphe/ChatGPT-4o-Writing-Prompts with responses regenerated with gemini-2.0-flash-thinking-exp-1219. Generation Details If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped. If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped. If ["candidates"][0]["finish_reason"] != 1 the sample was skipped. model =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/Gryphe_ChatGPT-4o-Writing-Prompts-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.texttext-generation1K<n<10K0 likes18 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.