datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
real-toxicity-prompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/allenai/real-toxicity-prompts.prompt-swap-mixed12-5xlr-e1-mxfp4-mergedprompt-swap-mixed12-5xlr-e2-mxfp4-mergedprompt-swap-medium12-e2-mxfp4-mergedllama-3.1-awesome-chatgpt-prompts
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/llama-3.1-awesome-chatgpt-prompts.prompt-safety-scores
Composite Safety Scoring for Prompts Using Multiple LLM Annotations
Introduction
Evaluating the safety of prompts is essential but challenging. Existing approaches often depend on predefined categories, which can be circumvented by new jailbreaks or attacks. Additionally, different tasks may require different safety thresholds.
This study explores using large language models (LLMs) themselves to annotate prompt safety. By combining these annotations, a continuous safety… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-safety-scores.cfg-sensitive-video-prompts
CFG-Sensitive Video Prompts
A provenance-first, text-only prompt suite for comparing classifier-free guidance behavior in video generation. It contains 22 English prompts:
existing_visualized: 6 prompts used by the local Wan2.2-TI2V-5B CFG-only LoRA scale-sweep report.
cfg_sensitive_candidate: 16 additional prompts selected because color/attribute binding, numeracy, readable text, spatial relations, temporal change, fluid motion, weather, camera motion, multi-subject… See the full description on the dataset page: https://huggingface.co/datasets/Perflow-Shuai/cfg-sensitive-video-prompts.kernelbench_with_promptsThis is a version of KernelBench where the prompts to produce the Triton and cuda kernel are explicitly saved in the JSON data files.
It only contains Level 1, 2, 3 kernels.
The prompt is the same as what is provided in the original KernelBench repo.
The dataset is prepared by Jiin Woo during her internship at AWS Annapurna Labs, the lab behind Trainium chips.
This dataset is part of an unreleased paper, and the paper will be updated in this README soon. If you use this dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/allenanie/kernelbench_with_prompts.prompt-swap-medium12-e1-mxfp4-mergedadaption-pokemon-story-prompts
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-pokemon_story_prompts
This dataset contains prompts instructing a model to write stories about specific Pokémon based on their detailed attributes, including stats, types, abilities, and lore. Each entry provides structured data such as height, weight, generation, and flavor text alongside an image URL. The primary focus is on generating creative narratives grounded… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/adaption-pokemon-story-prompts.webcode2m-scored-prompts-gpt-osswebcode2m-improved-promptsprompt-screening-dataset
Prompt Screening Dataset
This dataset is designed for training a classifier that identifies desirable rows for AI model training.
Each data source in agentlans/chatgpt contributes 10 000 rows.
Every row has been evaluated using these models:
agentlans/bge-small-en-v1.5-prompt-safety
agentlans/bge-small-en-v1.5-prompt-quality
agentlans/bge-small-en-v1.5-prompt-difficulty
agentlans/snowflake-arctic-embed-xs-refusal-classifierThe abridged column concatenates the input and output… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-screening-dataset.alpaca-prompts-annotated
Alpaca Annotated Dataset
This dataset includes prompts taken from yahma/alpaca-cleaned that have been annotated using the nvidia/prompt-task-and-complexity-classifier.
Each entry separates the instruction and input fields with two newline characters (\n\n).
The annotations describe the type of task and its complexity, as determined by NVIDIA’s classifier.
To know more about what each annotation means, see the classifier’s page on Hugging Face.
The prompts have been randomly… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/alpaca-prompts-annotated.adaption-hacker-news-article-prompts
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
hacker_news_article_prompts
This dataset contains text generation prompts derived from Hacker News article titles, instructing models to write full articles based on provided or self-generated headlines. The samples include specific tech and business topics from the Hacker News community, as well as instances where the model must invent a title from scratch. It is designed for… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/adaption-hacker-news-article-prompts.real-toxicity-prompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/krishna1707/real-toxicity-prompts.real-toxicity-prompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/ravi-softwarethreads/real-toxicity-prompts.tradehax-xai-grok-trading-visual-prompts
TradeHax xAI/Grok Trading Visual Prompts
Curated trading-scene image prompts and negative prompts tuned for xAI/Grok-inspired visual style.
Owner: Antired
Repo: tradehax-xai-grok-trading-visual-prompts
Synced by: scripts/sync-hf-assets.js
real-toxicity-prompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/wangdulou/real-toxicity-prompts.music-style-prompts
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
music_style_prompts
This dataset contains a collection of descriptive text prompts designed to generate diverse musical tracks across various genres, including pop, techno, ambient, and rock. Each entry details specific instrumentation, rhythmic patterns, atmospheric qualities, and emotional tones to guide audio synthesis. The content serves as a resource for training or… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/music-style-prompts.low-high-cost-promptsprompts_create_xprompt-sensitivity-codegen
Anonymous Prompt Sensitivity Dataset
This package contains model generations and evaluation outcomes for an anonymized
submission on prompt sensitivity in few-shot code generation.
What is included
prompt_sensitivity_dataset.jsonl: one row per generated sample
prompt_sensitivity_dataset.csv: tabular view of the same rows
prompt_sensitivity_dataset.parquet: columnar copy when parquet support is available
prompt_variant_spec.json: machine-readable description of the prompt… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-acl26/prompt-sensitivity-codegen.nation-parliamentary-promptsleo_prompts_v1
leo_prompts_v1: a dataset collection
collection of several prompts, nagative prompts and image urls datasets
data is uncleaned/non-normalized, they are as tey appear in leonardo.ai
data de-duplicated on a basic level.
contents
DatasetDict({
all: Dataset({
features: ['id', 'url', 'prompt', 'negative_prompt', 'imageHeight', 'imageWidth'],
num_rows: 299934
})
})
CITE
@misc {samact_2023,
author = { {SamAct} },
title =… See the full description on the dataset page: https://huggingface.co/datasets/SamAct/leo_prompts_v1.
