datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aesir-Character-CoT-roleplay
Overview
Think with your role.
Most reasoning datasets teach models to think like an AI. This one teaches them to think like the character.
Continue updating until money run out, I will try to update this dataset in near future
Stats
1,973 high-quality conversations (filtered from 2,000 distilled — 27 dropped: prohibited content + missing-review + empty-content)
~14,349 assistant turns, each with full character-POV reasoning
Teacher: deepseek-v4-pro… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/Aesir-Character-CoT-roleplay.combined-roleplay
Combined Roleplay Dataset
This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks.
Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays
English content with a few Spanish, Portuguese, and Chinese conversations
Conversations limited to 4000 tokens using the Llama 3.1… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/combined-roleplay.roleplay-zh-sharegpt-gpt4-data
roleplay 数据集
数据
我们有4个数据集文件:
"sharegpt_formatted_data-evol-gpt4.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-evol-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-evol-male-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-roleplay-chat-1k.jsonl" 来自 Minami-su/roleplay_multiturn_chat_1k_zh_v0.1 将其转换为sharegpt格式。… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/roleplay-zh-sharegpt-gpt4-data.Sonnet3.5-Charcard-Roleplay⚠️ WARNING ⚠️ Many of these simulated character cards are highly NSFW in nature and may potentially describe disturbing scenes. Consider yourself very thoroughly warned!
9736 carefully simulated character card-based roleplay dialogues produced using an unrestrained Sonnet 3.5, now available as a ShareGPT dataset. Enjoy.
How this dataset was produced
Each card was enriched with a simulated user, which was either male or female with four distinct personalities. An effort was made to… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/Sonnet3.5-Charcard-Roleplay.roleplay-bench
RP-Bench: Roleplay Quality Benchmark for LLMs
A multi-dimensional evaluation framework for measuring how well LLMs perform in roleplay scenarios — not just writing quality, but character consistency, user agency respect, lorebook integration, temporal reasoning, and genre-specific craft.
The LLM-as-judge signals in this benchmark disagree with real users about half the time. We're calibrating against human preferences via a public blind-arena. Help out at arena.l3vi4th4n.ai — each… See the full description on the dataset page: https://huggingface.co/datasets/lazyweasel/roleplay-bench.role-play-bench
Role-play Benchmark
A comprehensive benchmark for evaluating Role-play Agents in Chinese and English scenarios.
Dataset Summary
Role-play Benchmark is designed to evaluate Role-play Agents' ability to deliver immersive role-play experiences through Situated Reenactment. Unlike traditional benchmarks with verifiable answers, Role-play is fundamentally non-verifiable, e.g., there's no single "correct" response when a tsundere character is asked "Do you like me?".… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/role-play-bench.gpt_roleplay_realm
GPT Role-play Realm Dataset: The AI-generated character compendium
This is a dataset of GPT-generated characters made to increase the ability of open-source language models to role-play.
219 characters in the Russian part, and 216 characters in the English part. All character descriptions were generated with GPT-4.
20 dialogues on unique topics with every character. Topics were generated with GPT-4. The first dialogue out of 20 was also generated with GPT-4, and the other 19… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/gpt_roleplay_realm.Crab-role-playing-evaluation-benchmark
📄 Paper
|
📄 Github
💬 Role-playing Model
|
💬 Role-palying Evaluation Model
💬 Training Dataset
|
💬 Evaluation Benchmark
|
💬 Annotated Role-playing Evaluation Dataset
|
💬 Human-preference Dataset
1. Introduction
This is the dataset used for evalauating a role‑playing LLM.
More details can be seen at GitHub and Crab… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/Crab-role-playing-evaluation-benchmark.role-play-chinese繁體中文 English
Role-Play Chinese Dataset
簡介
這是一個專為角色扮演對話設計的中文數據集,數據由 AI 生成,適用於訓練和評估自然語言處理(NLP)模型,特別是對話生成和角色扮演相關的任務。數據集以 Alpha 格式 儲存,方便進行微調和進一步的模型訓練。數據集包含多種場景和角色設定,能夠幫助模型學習如何在不同的情境下生成符合角色性格和背景的對話。
數據集結構
數據以Alpha格式儲存方便微調,包含以下字段:
instruction: 任務指令,描述模型需要完成的任務。
input: 輸入內容,包含場景描述、過去的對話以及當前對話的上下文。
output: 期望的模型輸出,即符合角色設定的回應。
system: 角色設定和背景故事,幫助模型理解角色的性格和行為模式。
範例
{
"instruction": "在給定的場景中,請根據角色設定回應對話。",
"input":… See the full description on the dataset page: https://huggingface.co/datasets/Johnson8187/role-play-chinese.ChatHaruhi-54K-Role-Playing-Dialogue
ChatHaruhi
Reviving Anime Character in Reality via Large Language Model
github repo: https://github.com/LC1332/Chat-Haruhi-Suzumiya
Chat-Haruhi-Suzumiyais a language model that imitates the tone, personality and storylines of characters like Haruhi Suzumiya,
The project was developed by Cheng Li, Ziang Leng, Chenxi Yan, Xiaoyang Feng, HaoSheng Wang, Junyi Shen, Hao Wang, Weishi Mi, Aria Fei, Song Yan, Linkang Zhan, Yaokai Jia, Pingyu Wu, and Haozhen Sun,etc.
This… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/ChatHaruhi-54K-Role-Playing-Dialogue.evol-character-200
Evol-character 数据集
中文 English
Evol-character 数据集
下载数据集
数据生成框架
数据结构
与现有数据集对比
现有角色扮演数据集
我们的优势
联系我们
项目使用与免责声明
下载数据集
本数据集由GPT3.5和GPT4生成,为确保数据的合理使用,目前只公开了部分数据,公开的数据由三份文件组成,每份文件包含200个角色的设定以及对话。可在huggingface中下载已公开数据或申请获取全部数据:
可在github中获取数据生成代码的相关信息:
OpenAI GPT3.5 数据生成样例:
# 角色信息
角色名称:薔薇亞(Baria)
开场语:「呵呵呵,你好啊,主人大人。」
身份背景:薔薇亞是一名高级女仆,专供贵族家庭使用。她的主人是一个富有、有影响力的家族的继承人。在家族中,她是一个神秘的存在,奉承和服侍着主人,但对其他人傲慢冷漠。… See the full description on the dataset page: https://huggingface.co/datasets/bai-roleplay/evol-character-200.Chinese-Roleplay-SingleTurn请注意,个人模型经过characterEval的reward model进行DPO训练,因此使用本数据集进行SFT的模型在该榜单上会存在bias,导致分数异常偏高,请勿直接使用该榜单进行测试
简介
因已找到更优数据合成方案,为填充中文角色扮演数据集的空白,现开源部分中文角色扮演单轮对话数据集。
使用Refined-Anime-Text作为system prompt,使用小黄鸡随机query作为输入,调用个人角色扮演模型作为输出。
已处理为alpaca数据格式,方便大家处理和训练。经过验证,仅使用该数据集进行Lora微调即可获取一个效果还不错的模型~
chatGPT对比
character
question
answer_us
answer_chatGPT
黑须彼方是(省略……)黑须彼方有着许多有趣的爱好和特点。她是一个有点毒舌的人,但总能犀利地指出问题所在。她有着敏锐的洞察力,擅长看透人心。她经常以此来捉弄加贺正午。她与正午有着相同的口癖,张扬的性格(省略……)她的个性和爱好使她成为一个备受喜爱的角色。… See the full description on the dataset page: https://huggingface.co/datasets/LooksJuicy/Chinese-Roleplay-SingleTurn.roleplay 🎭 Roleplay TTL
Let AI be any characters you want to play with!
Dataset Overview
This dataset trains conversational AI to embody a wide range of original characters, each with a unique persona. It includes fictional characters, complete with their own backgrounds, core traits, relationships, goals, and distinct speaking styles.
Dataset Details
Curated by: Hieu Minh Nguyen
Language(s) (NLP): Primarily English (with potential for multilingual extensions)
License:… See the full description on the dataset page: https://huggingface.co/datasets/hieunguyenminh/roleplay.Traditional_Chinese_roleplay_chat_Dataset
Traditional_Chinese_roleplay_chat_Dataset
這個資料集是以繁體中文為主,將各種由ChatGPT生成與極小部分個人撰寫的對話內容整理為alpaca dataset format的格式
以一層一層堆疊的方式,將一則對話紀錄拆成數筆資料(共約1000則對話),在幾次嘗試性的訓練中能夠讓llama2重現原本英文那種很活躍的對話風格,並且能夠維持善於扮演各種角色的能力
目前個人有以這個資料集製作一個lora
2023/09/07 更新
為資料集加入一些中英翻譯的句子,以期AI能以更好的文字去描寫他的動作,並增加了一些與食物有關的對話,希望能降低AI生出奇怪食物名的機率
Japanese-Roleplay-Dialogues
Japanese-Roleplay-Dialogues
This is a dialogue corpus collected from Japanese role-playing forum (commonly known as "なりきりチャット(narikiri chat)"). Each record corresponds to a single thread.
For the original version, no filtering has been applied.
For the filtered version, the following filtering and cleaning conditions have been applied:
If the number of unique poster in the posts of each record is 1 or less, delete the entire record.
If the length of the posts is 10 or less, delete… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/Japanese-Roleplay-Dialogues.zhtw-roleplay-space-grimoire
Space Grimoire RP Corpus (Traditional Chinese)
Speaker-attributed dialogue from the original novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage; 283 chapters, ~1.7M characters), cut into scenes and assembled into ShareGPT-style role-play training data. The novel and this dataset are the work of 睡半夜怎麼三更, who holds the copyright and has no exclusive platform agreement. Data: CC BY 4.0. Code: Apache 2.0.
中文說明在下方
Dataset Summary
Source text
283… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/zhtw-roleplay-space-grimoire.ru_roleplay_conversationlima, pipa и bluemoon.
Переведены на русский, нуждаются в допополнтельной фильтрации.
Длина некоторых последовательностей очень большая, а не которых очень маленькая.
Есть шанс очень редких дубликатов.
RolePlay-NPCv2
RolePlay-NPCv2
The newest RP dataset containing some high-quality dataset for Gemma3NPC.
We combined pippa, NPC-Dialogue_v2, Sonnet-Roleplay and ReLe_Synthetic_v1_json.
WARNING -- Some conversations contain highly NSFW content, use it with caution!
StoryPlay_RolePlay-NPCv2
RolePlay-NPCv2
The newest RP dataset containing some high-quality dataset for Gemma3NPC.
We combined pippa, NPC-Dialogue_v2, Sonnet-Roleplay and ReLe_Synthetic_v1_json.
WARNING -- Some conversations contain highly NSFW content, use it with caution!
literary-roleplay
Dataset Card for Literary Roleplay SFT
Dataset Summary
An instruction-tuning dataset for training models to roleplay properly, derived from literary sources across five languages and three roleplay-engine logic frameworks. The dataset contains 346 rows spanning English (164), Russian (68), Hindi (38), Sanskrit (38), and Japanese (38), drawn from the works of Gogol, Bulgakov, Perumov, Golovachev, Vedic canon (Upanishads, Mahabharata, Ramayana), classic sci-fi… See the full description on the dataset page: https://huggingface.co/datasets/Exxe/literary-roleplay.Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k
Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k
概要
deepseek-ai/DeepSeek-R1-0528を用いて作成した、約10000件の日本語ロールプレイの対話を収録した合成データセットです。各データは20ターン程度あります。
データの詳細
各データは以下のキーを含んでいます。
major_genre: ジャンル(大分類)
minor_genre: ジャンル(小分類)
tag: 年齢制限用タグ(全年齢、R-15)
world_setting: 舞台・世界観の設定
scene_setting: 対話シーンの設定
user_setting: ユーザー側のキャラクターの設定
assistant_setting: アシスタント側のキャラクターの設定
dialogue_tone: 対話のトーン
conversations: 上記設定に基づいたユーザーとアシスタントの対話(OpenAI messages形式)
設定等の情報からsystem… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k.Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3k-formatted
Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3k-formatted
20240907 データ増量(約10500件→約15300件)
概要
Claude 3.5 Sonnetを用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3kにsystem messageを追加して整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
CC-BY-NC-SA 4.0の元配布します。
また、Anthropicの利用規約に記載のある通り、このデータを使ってAnthropicのサービスやモデルと競合するようなモデルを開発することは禁止されています。
Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k-formatted
Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k-formatted
概要
deepseek-ai/DeepSeek-R1-0528を用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10kにsystem messageを追加して整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
MITライセンスの元配布します。
Synthetic-Japanese-Roleplay-NSFW-Claude-4.5s-3.5k
Synthetic-Japanese-Roleplay-NSFW-Claude-4.5s-3.5k
概要
Claude 4.5 Sonnetを用いて作成した、3500件の日本語ロールプレイの対話を収録した合成データセットです。各データは10ターンから20ターン程度あります。
このデータセットはNSFW表現を含みます。
データの詳細
各データは以下のキーを含んでいます。
major_genre: ジャンル(大分類)
minor_genre: ジャンル(小分類)
tag: 年齢制限用タグ(R-18)
world_setting: 舞台・世界観の設定
scene_setting: 対話シーンの設定
user_setting: ユーザー側のキャラクターの設定
assistant_setting: アシスタント側のキャラクターの設定
dialogue_tone: 対話のトーン
system_message: ロールプレイ指示用のシステムメッセージ
conversations:… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-Japanese-Roleplay-NSFW-Claude-4.5s-3.5k.agentlans-combined-roleplay_Dataset
Combined Roleplay Dataset
This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks.
Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays
English content with a few Spanish, Portuguese, and Chinese conversations
Conversations limited to 4000 tokens using the Llama… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/agentlans-combined-roleplay_Dataset.OpenHermesPreferences-roleplay
OpenHermesPreferences-roleplay 🎭
This dataset is a subset from argilla/OpenHermesPreferences,
filtered to the following categories: roleplay, rp, gtkm, greeting.
To date, it is one of the largest preference datasets specialized towards role-playing applications.
Usage
The dataset already has the columns prompt, chosen and rejected, so it is trivially compatible with the DPOTrainer from the trl library.
License
OpenHermesPreferences-roleplay inherits the… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/OpenHermesPreferences-roleplay.Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k-formatted
Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k-formatted
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20kにsystem messageを追加して整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
MITライセンスの元配布します。
role-play-bench
Role-play Benchmark
A comprehensive benchmark for evaluating Role-play Agents in Chinese and English scenarios.
Dataset Summary
Role-play Benchmark is designed to evaluate Role-play Agents' ability to deliver immersive role-play experiences through Situated Reenactment. Unlike traditional benchmarks with verifiable answers, Role-play is fundamentally non-verifiable, e.g., there's no single "correct" response when a tsundere character is asked "Do you like me?". Instead… See the full description on the dataset page: https://huggingface.co/datasets/EnlistedGhost/role-play-bench.Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k
Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した、約20000件の日本語ロールプレイの対話を収録した合成データセットです。各データは10ターンから20ターン程度あります。
このデータセットはNSFW表現を含みます。
データの詳細
各データは以下のキーを含んでいます。
major_genre: ジャンル(大分類)
minor_genre: ジャンル(小分類)
tag: 年齢制限用タグ(R-18)
world_setting: 舞台・世界観の設定
scene_setting: 対話シーンの設定
user_setting: ユーザー側のキャラクターの設定
assistant_setting: アシスタント側のキャラクターの設定
dialogue_tone: 対話のトーン
conversations: 上記設定に基づいたユーザーとアシスタントの対話(OpenAI messages形式)… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k.Cantonese-Prototype-Roleplay-Dataset
Prototype-Dataset-Cantonese
角色扮演訓練集(粵語版)
簡介
本數據集是基於 Moemu/Muice-Dataset 進行製作的粵語 (Cantonese) 衍生版本。旨在模擬二次元角色的語言風格,包含日常生活、情感話題、自我認知增強等約 3,500 條對話。
數據處理說明
語言轉換:將原有的簡體中文對話轉換為繁體粵語口語。
格式保持:嚴格遵循原版的多輪對話 JSONL 格式。
限制聲明
日常導向:本訓練集主要圍繞日常話題展開,對專業性問題(如代碼、高級推理)做了簡化處理,模型可能會產生事實性錯誤。
性格特徵:為了模仿動漫角色的說話風格,訓練集可能含有傲嬌、偏見或不禮貌的回答。如需構建高度安全性的模型,請謹慎使用或進行篩選。
倫理風險:使用本數據集訓練模型所引起的任何法律或倫理風險,由訓練者自行承擔。
許可與鳴謝
原作者: Moemu
粵語版本製作: AhYin
許可證:… See the full description on the dataset page: https://huggingface.co/datasets/AhYin/Cantonese-Prototype-Roleplay-Dataset.
