datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k
Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k
概要
deepseek-ai/DeepSeek-R1-0528を用いて作成した、約10000件の日本語ロールプレイの対話を収録した合成データセットです。各データは20ターン程度あります。
データの詳細
各データは以下のキーを含んでいます。
major_genre: ジャンル(大分類)
minor_genre: ジャンル(小分類)
tag: 年齢制限用タグ(全年齢、R-15)
world_setting: 舞台・世界観の設定
scene_setting: 対話シーンの設定
user_setting: ユーザー側のキャラクターの設定
assistant_setting: アシスタント側のキャラクターの設定
dialogue_tone: 対話のトーン
conversations: 上記設定に基づいたユーザーとアシスタントの対話(OpenAI messages形式)
設定等の情報からsystem… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k.Japanese-RP-Bench-testdata-SFW
Japanese-RP-Bench-testdata-SFW
本データセットは、LLMの日本語ロールプレイ能力を計測するベンチマークJapanese-RP-Bench用の評価データセットです。
ベンチマークの詳細については記事を参照してください。
データの概要
本データは以下のようなキーを持ちます。
genre: ロールプレイのジャンル
tag: ロールプレイの年齢区分
world_setting: ロールプレイの世界観設定
scene_setting: ロールプレイのシーン設定
user_setting: ロールプレイのユーザー側キャラクター設定
assistant_setting: ロールプレイのアシスタント側キャラクター設定
dialogue_tone: ロールプレイの対話のトーン
first_user_input: ロールプレイの最初のユーザー発話
response_format: ロールプレイの応答形式
id: データのid
ライセンス… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Japanese-RP-Bench-testdata-SFW.Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k-formatted
Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k-formatted
概要
deepseek-ai/DeepSeek-R1-0528を用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10kにsystem messageを追加して整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
MITライセンスの元配布します。
genre-taxonomy-sfw
Genre Taxonomy — SFW Split
The safe-for-work half of a two-part short-fiction genre taxonomy: 119,753 genre
entries / 197,509 labels (genre + subgenre names; no story text). The adult
companion split lives in the paired repo genre-taxonomy-nsfw.
Every label was screened by multi-round LLM judging (large single-pass scan, then
targeted re-judging rounds) under a double-pass agreement standard: a label is
only acted on when independent judging passes agree. Labels judged… See the full description on the dataset page: https://huggingface.co/datasets/baiango/genre-taxonomy-sfw.Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k-formatted
Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k-formatted
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20kにsystem messageを追加して整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
MITライセンスの元配布します。
Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k
Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した、約20000件の日本語ロールプレイの対話を収録した合成データセットです。各データは10ターンから20ターン程度あります。
データの詳細
各データは以下のキーを含んでいます。
major_genre: ジャンル(大分類)
minor_genre: ジャンル(小分類)
tag: 年齢制限用タグ(全年齢、R-15)
world_setting: 舞台・世界観の設定
scene_setting: 対話シーンの設定
user_setting: ユーザー側のキャラクターの設定
assistant_setting: アシスタント側のキャラクターの設定
dialogue_tone: 対話のトーン
conversations: 上記設定に基づいたユーザーとアシスタントの対話(OpenAI messages形式)
設定等の情報からsystem… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k.Reddit-SFW-Writing_Prompts_ShareGPT_Curated
Normalized SFW Reddit Writing Prompts
Dataset Description
This dataset is a normalized, flattened version of curated Reddit writing prompts, specifically derived from ChaoticNeutrals/Reddit-SFW-Writing_Prompts_ShareGPT. It maps nested conversational arrays into a strict instruction-response schema, making it highly optimized for instruction-tuning Large Language Models.
Dataset Schema
Column Name
Type
Description
prompt
string
The input prompt, user… See the full description on the dataset page: https://huggingface.co/datasets/rafy2342/Reddit-SFW-Writing_Prompts_ShareGPT_Curated.roleplay-conversations-arabic-sfw-nsfw
Roleplay Conversations Arabic SFW-NSFW
A multilingual collection of 914 roleplay conversations for training conversational AI models.
Dataset Structure
Each entry contains:
messages: Array of conversation turns with role (system/user/assistant) and content
Languages
Arabic (Gulf & Egyptian dialects) - 85%
English - 3%
French - 2%
Japanese - 2%
Categories
SFW: Safe-for-work roleplay conversations
NSFW: Adult roleplay conversations… See the full description on the dataset page: https://huggingface.co/datasets/TYDTYDYT/roleplay-conversations-arabic-sfw-nsfw.
