datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aesir-Character-CoT-roleplay
Overview
Think with your role.
Most reasoning datasets teach models to think like an AI. This one teaches them to think like the character.
Continue updating until money run out, I will try to update this dataset in near future
Stats
1,973 high-quality conversations (filtered from 2,000 distilled — 27 dropped: prohibited content + missing-review + empty-content)
~14,349 assistant turns, each with full character-POV reasoning
Teacher: deepseek-v4-pro… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/Aesir-Character-CoT-roleplay.japanese-young-character-roleplay-vlm
Japanese Young Character Roleplay VLM
日本語のVLM向けに作られた、画像接地型ロールプレイSFTデータセットです。幼い雰囲気の
完全な架空キャラクターが、自分を名前で呼びながら、画像について安全な日常会話を
続けます。
既存作品のキャラクター、実在人物、Webから取得した画像は一切含みません。画像・会話・
メタデータは、このリポジトリの決定的な手続き生成器だけで作成しています。
規模
split
rows
assistant turns
train
98,000
294,000
validation
1,000
3,000
test
1,000
3,000
total
100,000
300,000
各行には、256×192 WebP画像が1枚と、systemを含む7メッセージ
(user/assistant 3往復)が入っています。48の架空名、12の場面、16の物体、8色、
6タスクを決定的に組み合わせています。… See the full description on the dataset page: https://huggingface.co/datasets/beezza/japanese-young-character-roleplay-vlm.deepfabric-character-roleplaykorean-character-roleplay-sft
Korean Character Roleplay SFT Dataset
Character-based Korean roleplay conversation dataset for fine-tuning language models.
Dataset Description
This dataset contains high-quality Korean roleplay conversations between users and AI characters. Each conversation follows a specific character's personality, speech patterns, and voice profile.
Dataset Statistics
Split
Samples
Train
965
Test
108
Total
1,073
Quality Metrics
Overall… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/korean-character-roleplay-sft.evol-character-entire
Evol-character 数据集
中文 English
Evol-character 数据集
下载数据集
数据生成框架
数据结构
与现有数据集对比
现有角色扮演数据集
我们的优势
联系我们
项目使用与免责声明
下载数据集
本数据集由GPT3.5和GPT4生成,为确保数据的合理使用,目前只公开了部分数据,公开的数据由三份文件组成,每份文件包含200个角色的设定以及对话。可在huggingface中下载已公开数据或申请获取全部数据:
可在github中获取数据生成代码的相关信息:
OpenAI GPT3.5 数据生成样例:
# 角色信息
角色名称:薔薇亞(Baria)
开场语:「呵呵呵,你好啊,主人大人。」
身份背景:薔薇亞是一名高级女仆,专供贵族家庭使用。她的主人是一个富有、有影响力的家族的继承人。在家族中,她是一个神秘的存在,奉承和服侍着主人,但对其他人傲慢冷漠。… See the full description on the dataset page: https://huggingface.co/datasets/bai-roleplay/evol-character-entire.character-roleplay-DPOThis is practical-dreamer/RPGPT_PublicDomain-alpaca with an added rejection column generated by Phi3-q4.
flock-off-s1-character-roleplay
