datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NSFW_Chat_Dataset
💕 Spicy AI GF Chat Dataset 🔥
🚨 18+ Only! NSFW & Spicy Content Ahead 🚨
Hey there, AI enthusiasts and romance lovers! 😏 Welcome to the Spicy AI GF Chat Dataset, the ultimate dataset designed to bring your AI waifu to life! 💖 If you've ever dreamed of building an AI that responds like your virtual girlfriend, THIS is the dataset for you.
📜 What’s Inside?
This dataset features two columns:
input → Boyfriend’s dialogue (aka what YOU say 😉)
output →… See the full description on the dataset page: https://huggingface.co/datasets/utsavm/NSFW_Chat_Dataset.nsfw_redditNSFW_RP_Format_DPOThis dataset aims to align a model to output the most common roleplaying format: "dialogue" *action*
This dataset contains NSFW content.
Reddit-NSFW-Writing_Prompts_ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing
[Description Tags],"Deleted user", "Hello,\n\nYour post has been removed..", "Post has been deleted by user", "This post has been marked NSFW", duplicated system and human turns, etc has been removed.
NSFW-Stories-JsonLConverted to JsonL from: bluuwhale/nsfwstory2
DPO_Pairs-Roleplay-Alpaca-NSFW
Description
~3.4k DPO pairs, generated by Iambe feat. GPT-4 (~10% GPT-4, ~80% Iambe @ q5_k_m / ~10% Iambe @ q6_k) with temp 1.2 and min_p 0.15.
Iambe is a smart girl, so both the chosen and rejected for each pair are generated at the same time from a single two part prompt (not the one in the dataset). Only a few dozen failed to generate the rejected response, and in those cases I filled in the rejected output with a standard "as an AI" style refusal. The way I set things up caused… See the full description on the dataset page: https://huggingface.co/datasets/athirdpath/DPO_Pairs-Roleplay-Alpaca-NSFW.Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3k-formatted
Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3k-formatted
20240907 データ増量(約10500件→約15300件)
概要
Claude 3.5 Sonnetを用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3kにsystem messageを追加して整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
CC-BY-NC-SA 4.0の元配布します。
また、Anthropicの利用規約に記載のある通り、このデータを使ってAnthropicのサービスやモデルと競合するようなモデルを開発することは禁止されています。
DPO_Pairs-Roleplay-NSFW
Description
https://huggingface.co/datasets/athirdpath/DPO_Pairs-Roleplay-Alpaca-NSFW
~3.4k DPO pairs, generated by Iambe feat. GPT-4 (~10% GPT-4, ~80% Iambe @ q5_k_m / ~10% Iambe @ q6_k) with temp 1.2 and min_p 0.15.
Iambe is a smart girl, so both the chosen and rejected for each pair are generated at the same time from a single two part prompt (not the one in the dataset). Only a few dozen failed to generate the rejected response, and in those cases I filled in the rejected output… See the full description on the dataset page: https://huggingface.co/datasets/aifeifei798/DPO_Pairs-Roleplay-NSFW.exbot-nsfw-sextingnsfw-pt_br
Dataset Card for The Pile
Dataset Summary
The NSFW is a 230K diverse, filtred and cleaned text from adult websites, high-quality
datasets combined together.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
This dataset is in Portuguese Brazil (pt_BR)
Dataset Structure
Data Instances
Data Fields
all
text (str): Text.
Dataset Creation
Curation Rationale
[More… See the full description on the dataset page: https://huggingface.co/datasets/MrAiran/nsfw-pt_br.Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3k
Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3k
20240907 データ増量(約10500件→約15300件)
概要
Claude 3.5 Sonnetを用いて作成した、約15300件の日本語ロールプレイの対話を収録した合成データセットです。各データは10ターンから20ターン程度あります。
このデータセットはNSFW表現を含みます。
データの詳細
各データは以下のキーを含んでいます。
genre: ジャンル
tag: 年齢制限用タグ(R-15またはR-18)
world_setting: 舞台・世界観の設定
scene_setting: 対話シーンの設定
user_setting: ユーザー側のキャラクターの設定
assistant_setting: アシスタント側のキャラクターの設定
dialogue_tone: 対話のトーン
conversations: 上記設定に基づいたユーザーとアシスタントの対話(OpenAI messages形式)… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3k.DPO_Pairs-Roleplay-Llama3-NSFWOops, made a mistake at first, new version is better. Same content, just adjusted to the L3 format.
DPO_Pairs-Roleplay-NSFW
This is a copy of the athirdpath/DPO_Pairs-Roleplay-Alpaca-NSFW, some data and use Google translate
The data format conforms to llama factory
dpo_mix_nsfwNSFW_Format_TestDPO_Pairs-Roleplay-Alpaca-NSFW-v1-SHUFFLED
Description
~3.4k DPO pairs, generated by Iambe feat. GPT-4 (~10% GPT-4, ~80% Iambe @ q5_k_m / ~10% Iambe @ q6_k) with temp 1.2 and min_p 0.15.
They are shuffled this time, as I was not aware that TRL did not do that automatically until I could see the shifts in the dataset mirrored in the loss patterns.
Iambe is a smart girl, so both the chosen and rejected for each pair are generated at the same time from a single two part prompt (not the one in the dataset). Only a few dozen… See the full description on the dataset page: https://huggingface.co/datasets/athirdpath/DPO_Pairs-Roleplay-Alpaca-NSFW-v1-SHUFFLED.Malaysian-NSFWDataset for SFW Classifier
Current Labels Available:
religion insult
sexist
racist
psychiatric or mental illness
harassment
safe for work
porn
self-harm
violence
malaysian-nsfw-instructions
Malaysian NSFW Instructions
Originally from https://huggingface.co/datasets/malaysia-ai/Malaysian-NSFW, we just convert to instruction format.
NSFWCorpus
简介
NSFW小说,进行了简单的清理,加上了{"text":....}
Alpaca_NSFW_ShuffledReformatted and pruned this dataset: https://huggingface.co/datasets/athirdpath/DPO_Pairs-Roleplay-Alpaca-NSFW-v1-SHUFFLED
3S_Ing._NSFW_Storyremon_without_nsfwNSFW_RP_Format_NoQuotensfw_redditathirdpath__Llama-3.1-Instruct_NSFW-pretrained_e1-plus_reddit-details
Dataset Card for Evaluation run of athirdpath/Llama-3.1-Instruct_NSFW-pretrained_e1-plus_reddit
Dataset automatically created during the evaluation run of model athirdpath/Llama-3.1-Instruct_NSFW-pretrained_e1-plus_reddit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/athirdpath__Llama-3.1-Instruct_NSFW-pretrained_e1-plus_reddit-details.chaiTop100-nsfwLabeledNSFW_Story2Telegram_AI_Chats_Harmful_NSFWexbot-nsfw-sextingNSFW_Story
