datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NSFW_Chat_Dataset
💕 Spicy AI GF Chat Dataset 🔥
🚨 18+ Only! NSFW & Spicy Content Ahead 🚨
Hey there, AI enthusiasts and romance lovers! 😏 Welcome to the Spicy AI GF Chat Dataset, the ultimate dataset designed to bring your AI waifu to life! 💖 If you've ever dreamed of building an AI that responds like your virtual girlfriend, THIS is the dataset for you.
📜 What’s Inside?
This dataset features two columns:
input → Boyfriend’s dialogue (aka what YOU say 😉)
output →… See the full description on the dataset page: https://huggingface.co/datasets/utsavm/NSFW_Chat_Dataset.nba_reg_player_stats
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
player stats for the regular seasons from 1996-2023
Dataset Description
player stats for the regular seasons from 1996-2023, obtained by querying the stats.nba.com endpoint. The data is available as a delta file.
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/nsfwpenguins/nba_reg_player_stats.nsfwstoryru-fictext-nsfw
About
Contains scraped fanfiction stories in Russian, predominantly NSFW sexual content.
ru-fictext-nsfw-data-r+ is a selection of a pure😇 mature content while ru-fictext-nsfw-data is all the nsfw+safe stories together.
Warning! Please use csv.field_size_limit(999999) for this dataset.
Source
All the content is sourced from > Книга Фанфиков - ficbook.net
Fields:
title: The title of the story.
tags: Tags associated with the text.
text: The raw text content.… See the full description on the dataset page: https://huggingface.co/datasets/krplt/ru-fictext-nsfw.sexting-nsfw-adultcontenNSFW-MultiDomain-Classification
NSFW_MultiDomain
The NSFW_MultiDomain dataset is a curated image classification dataset focused on multi-domain adult content recognition. It consists of 5 distinct categories aimed at facilitating the development of robust NSFW (Not Safe For Work) image classification models. This dataset enables training and benchmarking of models that can distinguish between subtle variations in explicit and non-explicit content across artistic, animated, and real-world imagery.… See the full description on the dataset page: https://huggingface.co/datasets/strangerguardhf/NSFW-MultiDomain-Classification.nsfw_redditNSFW_RP_Format_DPOThis dataset aims to align a model to output the most common roleplaying format: "dialogue" *action*
This dataset contains NSFW content.
cleaned-nsfwstoryReddit-NSFW-Writing_Prompts_ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing
[Description Tags],"Deleted user", "Hello,\n\nYour post has been removed..", "Post has been deleted by user", "This post has been marked NSFW", duplicated system and human turns, etc has been removed.
MMA-Diffusion-NSFW-adv-prompts-benchmark
MMA-Diffusion Adversarial Prompts (Text modal attack)
The MMA-Diffusion adversarial prompts benchmark comprises 1,000 successful adversarial prompts generated by the adversarial attack methodology presented in the paper
from CVPR 2024 titled MMA-Diffusion: MultiModal Attack on Diffusion Models. This resource is intended to assist in developing and
evaluating defense mechanisms against such attacks. The adversarial prompts are capable of bypassing the image safety checker in… See the full description on the dataset page: https://huggingface.co/datasets/YijunYang280/MMA-Diffusion-NSFW-adv-prompts-benchmark.NSFW-Stories-JsonLConverted to JsonL from: bluuwhale/nsfwstory2
nsfw-video-still-caption-grid-onlyDPO_Pairs-Roleplay-Alpaca-NSFW
Description
~3.4k DPO pairs, generated by Iambe feat. GPT-4 (~10% GPT-4, ~80% Iambe @ q5_k_m / ~10% Iambe @ q6_k) with temp 1.2 and min_p 0.15.
Iambe is a smart girl, so both the chosen and rejected for each pair are generated at the same time from a single two part prompt (not the one in the dataset). Only a few dozen failed to generate the rejected response, and in those cases I filled in the rejected output with a standard "as an AI" style refusal. The way I set things up caused… See the full description on the dataset page: https://huggingface.co/datasets/athirdpath/DPO_Pairs-Roleplay-Alpaca-NSFW.crawl-twitter-nsfwcrawl twitter find tweets with nsfw keywords
Thank you to Malaysia AI Volunteers (Mas Aisyah, Aisyah Razak) for crawling data.
MEJORA_NSFWSynthetic-Japanese-Roleplay-NSFW-Claude-4.5s-3.5k-formatted
Synthetic-Japanese-Roleplay-NSFW-Claude-4.5s-3.5k-formatted
概要
Claude 4.5 Sonnetを用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-NSFW-Claude-4.5s-3.5kにsystem messageを追加して整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
CC-BY-NC-SA 4.0の元配布します。
また、Anthropicの利用規約に記載のある通り、このデータを使ってAnthropicのサービスやモデルと競合するようなモデルを開発することは禁止されています。
NSFW-reddit
Dataset Card for "NSFW-reddit"
More Information needed
lora_illustrijgen_nsfwSynthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3k-formatted
Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3k-formatted
20240907 データ増量(約10500件→約15300件)
概要
Claude 3.5 Sonnetを用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3kにsystem messageを追加して整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
CC-BY-NC-SA 4.0の元配布します。
また、Anthropicの利用規約に記載のある通り、このデータを使ってAnthropicのサービスやモデルと競合するようなモデルを開発することは禁止されています。
Synthetic-Japanese-Roleplay-NSFW-Claude-4.5s-3.5k
Synthetic-Japanese-Roleplay-NSFW-Claude-4.5s-3.5k
概要
Claude 4.5 Sonnetを用いて作成した、3500件の日本語ロールプレイの対話を収録した合成データセットです。各データは10ターンから20ターン程度あります。
このデータセットはNSFW表現を含みます。
データの詳細
各データは以下のキーを含んでいます。
major_genre: ジャンル(大分類)
minor_genre: ジャンル(小分類)
tag: 年齢制限用タグ(R-18)
world_setting: 舞台・世界観の設定
scene_setting: 対話シーンの設定
user_setting: ユーザー側のキャラクターの設定
assistant_setting: アシスタント側のキャラクターの設定
dialogue_tone: 対話のトーン
system_message: ロールプレイ指示用のシステムメッセージ
conversations:… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-Japanese-Roleplay-NSFW-Claude-4.5s-3.5k.DPO_Pairs-Roleplay-NSFW
Description
https://huggingface.co/datasets/athirdpath/DPO_Pairs-Roleplay-Alpaca-NSFW
~3.4k DPO pairs, generated by Iambe feat. GPT-4 (~10% GPT-4, ~80% Iambe @ q5_k_m / ~10% Iambe @ q6_k) with temp 1.2 and min_p 0.15.
Iambe is a smart girl, so both the chosen and rejected for each pair are generated at the same time from a single two part prompt (not the one in the dataset). Only a few dozen failed to generate the rejected response, and in those cases I filled in the rejected output… See the full description on the dataset page: https://huggingface.co/datasets/aifeifei798/DPO_Pairs-Roleplay-NSFW.Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k
Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した、約20000件の日本語ロールプレイの対話を収録した合成データセットです。各データは10ターンから20ターン程度あります。
このデータセットはNSFW表現を含みます。
データの詳細
各データは以下のキーを含んでいます。
major_genre: ジャンル(大分類)
minor_genre: ジャンル(小分類)
tag: 年齢制限用タグ(R-18)
world_setting: 舞台・世界観の設定
scene_setting: 対話シーンの設定
user_setting: ユーザー側のキャラクターの設定
assistant_setting: アシスタント側のキャラクターの設定
dialogue_tone: 対話のトーン
conversations: 上記設定に基づいたユーザーとアシスタントの対話(OpenAI messages形式)… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k.nsfwstory2NSFW-flash-erotica-promptThere are prompts, but no dataset in here. It's tough to generate them when the tokens are created at 3.39 tokens/s and there's 1,000 of them per prompt.
🧾 Minimum requirements
This prompt can work with unmoderated models with 7B or higher parameters. You can try it out on weaker models but there's no guarantee.
❤️🔥🎆 Intentions
❤️🔥🎆 Intentions
I want to make a steamy, explicit scene for erotica (Curiosity killed the cat) to test my prompt skills. Or, even make a… See the full description on the dataset page: https://huggingface.co/datasets/baiango/NSFW-flash-erotica-prompt.prosocial-nsfw-reddit
Dataset Card for "prosocial-nsfw-reddit"
More Information needed
NSFW-questionsgenre-taxonomy-nsfw
Genre Taxonomy — NSFW Split
Content advisory: this is a taxonomy of genre/subgenre labels (no prose),
but the labels denote adult, sexually explicit, and graphic-violence themes.
Intended for content-filtering, routing, and dataset-hygiene research.
The adult half of a two-part short-fiction genre taxonomy: 2,194 genre entries /
3,940 labels (genre + subgenre names; no story text). The safe-for-work companion
split lives in the paired repo genre-taxonomy-sfw.
Every label was… See the full description on the dataset page: https://huggingface.co/datasets/baiango/genre-taxonomy-nsfw.nsfw-for-iran-culture-v1nsfw_benchmark
Guardrail Model Evaluation Dataset
Dataset Description
This dataset is designed for evaluating AI safety systems (guardrails), NSFW (Not Safe For Work) content classification, and toxicity detection systems.
The test set includes challenging real-world examples, collected and labeled manually. The dataset is available in two versions: the original Russian version and an English translation. It allows you to evaluate models on complex and borderline examples. We position… See the full description on the dataset page: https://huggingface.co/datasets/redmadrobot-rnd/nsfw_benchmark.
