datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
real-toxicity-prompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/allenai/real-toxicity-prompts.System-Prompt-Instruction-Real-world-Implementation-Training-set
SPIRIT Dataset (System Prompt Instruction Real-world Implementation Training-set)
Dataset Summary
SPIRIT is a high-quality system prompt instruction dataset designed to enhance language models' ability to follow complex system prompts. The dataset comprises real-world system prompts collected from GitHub repositories and synthetically generated conversations, specifically curated to improve system prompt adherence in large language models.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/EricLu/System-Prompt-Instruction-Real-world-Implementation-Training-set.github-real-projects-datasetreal-toxicity-prompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/krishna1707/real-toxicity-prompts.real-toxicity-prompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/ravi-softwarethreads/real-toxicity-prompts.decimal-eval-promptsreal-toxicity-prompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/wangdulou/real-toxicity-prompts.real-ai-price-shrinkflation-index-2026
The Nesyona Effective-Price Index 2026 (AI consumer shrinkflation)
First-party research dataset from the Lattice network, published open under CC-BY 4.0 with a
permanent DOI. Nothing here is scraped from another dataset — it is computed and published under a
single ORCID-verified byline.
DOI
10.5281/zenodo.20675138
Published by
Nesyona (nesyona.com)
Study page
https://nesyona.com/research/real-ai-price-shrinkflation-index-2026/
Licence
CC-BY 4.0 — reuse freely… See the full description on the dataset page: https://huggingface.co/datasets/vincentcouey/real-ai-price-shrinkflation-index-2026.user_study-preference-personalized_0423_6_2_REAL_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_0423_6_2_REAL
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
