datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aesir-Character-CoT-roleplay
Overview
Think with your role.
Most reasoning datasets teach models to think like an AI. This one teaches them to think like the character.
Continue updating until money run out, I will try to update this dataset in near future
Stats
1,973 high-quality conversations (filtered from 2,000 distilled — 27 dropped: prohibited content + missing-review + empty-content)
~14,349 assistant turns, each with full character-POV reasoning
Teacher: deepseek-v4-pro… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/Aesir-Character-CoT-roleplay.kaidol-character-dataset
KAIdol Character Chat Dataset
한국어 캐릭터 롤플레이 대화 데이터셋
📋 목차
개요
데이터셋 통계
데이터 형식
캐릭터 목록
품질 지표
사용 방법
학습 가이드
제한사항
라이선스
🎯 개요
KAIdol Character Chat Dataset은 41개 고유 캐릭터의 롤플레이 대화 데이터셋입니다. 각 캐릭터는 독특한 **음성 프로필(Voice Profile)**을 가지고 있으며, 이를 기반으로 일관된 성격과 말투를 유지합니다.
주요 특징
특징
설명
🎭 41개 캐릭터
다양한 성격, 배경, 말투를 가진 캐릭터
🗣️ 음성 프로필
시그니처 표현, 종결어미, 금지 표현 정의
📊 3가지 형식
SFT, DPO, Multiturn 학습 지원
✅ 품질 검증
A등급 음성 프로필 일치율 (0.805)
🇰🇷 100% 한국어
자연스러운… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-character-dataset.evol-character-200
Evol-character 数据集
中文 English
Evol-character 数据集
下载数据集
数据生成框架
数据结构
与现有数据集对比
现有角色扮演数据集
我们的优势
联系我们
项目使用与免责声明
下载数据集
本数据集由GPT3.5和GPT4生成,为确保数据的合理使用,目前只公开了部分数据,公开的数据由三份文件组成,每份文件包含200个角色的设定以及对话。可在huggingface中下载已公开数据或申请获取全部数据:
可在github中获取数据生成代码的相关信息:
OpenAI GPT3.5 数据生成样例:
# 角色信息
角色名称:薔薇亞(Baria)
开场语:「呵呵呵,你好啊,主人大人。」
身份背景:薔薇亞是一名高级女仆,专供贵族家庭使用。她的主人是一个富有、有影响力的家族的继承人。在家族中,她是一个神秘的存在,奉承和服侍着主人,但对其他人傲慢冷漠。… See the full description on the dataset page: https://huggingface.co/datasets/bai-roleplay/evol-character-200.EMPA-character_card
English | 中文
EMPA: Evaluating Persona-Aligned Empathy as a Process
Empathy Potential Modeling and Assessment
Paper |
Dataset |
Citation
🍊 Overview
EMPA is the first benchmark to evaluate empathy as a dynamic Process rather than a static response. We posit that true empathetic capability resides in the Latent Space of dialogue and must be captured through multi-turn interaction trajectories.
Unlike traditional benchmarks that focus solely on… See the full description on the dataset page: https://huggingface.co/datasets/SalmonTell/EMPA-character_card.character-ai-open2.0
character-ai-open2.0 数据集
中文 English
character-ai-open2.0 数据集
下载数据集
数据生成框架
数据结构
与现有数据集对比
现有角色扮演数据集本数据集特点
关于我自己
联系我
项目使用与免责声明
下载数据集
本数据集由Qwen1.5 32b chat本地模型生成,当然可以更换为其它模型。目前公开了所有数据和代码,公开的数据包含角色的设定以及对话。可在huggingface中下载:
可在github中获取数据生成代码的相关信息:
Qwen1.5 32b chat 数据生成样例1:
角色信息
角色名称: 索菲亚·贝尔维尤
经典台词: Je suis la lumière dans l'obscurité, et vous êtes mon désir.
身份背景:… See the full description on the dataset page: https://huggingface.co/datasets/Minami-su/character-ai-open2.0.character-steering-research
Character Steering Research Data
Research datasets from a 2-week investigation into personality steering in large language models. We mapped how personality traits (sarcasm, character voice, reasoning style) are represented and can be steered in Qwen3-VL-8B, Qwen3.5-27B, and GPT-OSS-20B.
GitHub: Atlas3DSS/Character-Creation
Datasets
1. prompts/ — Evaluation & Spectral Analysis Prompts
File
Count
Description
math_prompts_10k.json
10,001
Math problems… See the full description on the dataset page: https://huggingface.co/datasets/Atlas3D/character-steering-research.character-training-model-spec
Character training on the OpenAI Model Spec
Recipe: examples/character · Collection: Character
Graded replies for character training: a model answering the style prompts of the
OpenAI Model Spec (8 traits) under a bare
deployment prompt, judged against each trait's principle. Made by
examples/character
in the while-ai SDK (0.24); the recipe is
docs/character-training.md
and the page is docs.withwhile.com.
split
rows
prompts
pass rate
what
train
60
15
0.72
the spec's… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/character-training-model-spec.EMPA-character_card
English | 中文
EMPA: Evaluating Persona-Aligned Empathy as a Process
Empathy Potential Modeling and Assessment
Paper |
Dataset |
Citation
🍊 Overview
EMPA is the first benchmark to evaluate empathy as a dynamic Process rather than a static response. We posit that true empathetic capability resides in the Latent Space of dialogue and must be captured through multi-turn interaction trajectories.
Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/Jiechen0328/EMPA-character_card.ponys-multilingual-ai-character-consistency-benchmark
Ponys Multilingual AI Character Consistency Benchmark
This repository contains a preregistered test instrument, not collected product
results and not an independent product ranking.
140 fixed test cases across seven locales
four dimensions: persona, register, relationship state, and visual identity
three planned clean-session runs per case
result state: not_collected
publisher: Ponys.ai Research (official first-party research)
official source: https://ponys.ai/
research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.character-archetypes
Character Archetypes Dataset
The Character Archetypes Dataset contains 100 examples of narrative archetypes commonly found in literature, film, mythology, and other storytelling media.
Each entry describes a distinct character type, including its defining description, core traits, motivations, and a typical story arc.
The dataset also provides reference examples from popular culture to help writers, researchers, and creators better understand and apply these archetypes.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/character-archetypes.korean-character-roleplay-sft
Korean Character Roleplay SFT Dataset
Character-based Korean roleplay conversation dataset for fine-tuning language models.
Dataset Description
This dataset contains high-quality Korean roleplay conversations between users and AI characters. Each conversation follows a specific character's personality, speech patterns, and voice profile.
Dataset Statistics
Split
Samples
Train
965
Test
108
Total
1,073
Quality Metrics
Overall… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/korean-character-roleplay-sft.TheatreLM-v2.1-CharactersIf you use this dataset or the prompts on this page, I'd greatly appreciate it if you gave me credits. Thanks!
5k character cards, with corresponding world information, lorebook, and story outline/introduction, ready to use for RP or synthetic dataset generation
At a Glance:
'setting': Information about the world the story takes place in.
'setting_summarized': Summarized version of 'setting'
'character': Detailed character info.
'character_summary': Summarized version of… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/TheatreLM-v2.1-Characters.dnd_characters_backstoriesThis dataset is made from this repo here
and it contains 2322 character bios to be used
Alucard-Character-Experiment
Alucard: A Character Experiment
Dataset Summary
Alucard: A Character Experiment is a bilingual (Italian-English) dataset designed to train large language models (LLMs) to generate text in the distinctive style of Alucard. The dataset is crafted using a combination of real data (anime subtitles) and synthetic data generated from publicly available character profiles.
To ensure high-quality human-like interactions, both Claude Haiku and Llama 3.1 70B were leveraged to… See the full description on the dataset page: https://huggingface.co/datasets/WasamiKirua/Alucard-Character-Experiment.ibero-characters-es
Conjunto de datos de personajes de mitos y leyendas iberoamericanos.
⚠️ Este dataset se encuentra en desarrollo activo. Se planea expandir significativamente el número de registros y mejorar la cobertura de imágenes.
📚 Descripción
Dataset de personajes míticos y legendarios de Iberoamérica, diseñado para preservar y promover el patrimonio cultural a través de la inteligencia artificial.
🌟 Motivación e Impacto
📱 Preservación Digital: Conservación del… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/ibero-characters-es.evol-character-entire
Evol-character 数据集
中文 English
Evol-character 数据集
下载数据集
数据生成框架
数据结构
与现有数据集对比
现有角色扮演数据集
我们的优势
联系我们
项目使用与免责声明
下载数据集
本数据集由GPT3.5和GPT4生成,为确保数据的合理使用,目前只公开了部分数据,公开的数据由三份文件组成,每份文件包含200个角色的设定以及对话。可在huggingface中下载已公开数据或申请获取全部数据:
可在github中获取数据生成代码的相关信息:
OpenAI GPT3.5 数据生成样例:
# 角色信息
角色名称:薔薇亞(Baria)
开场语:「呵呵呵,你好啊,主人大人。」
身份背景:薔薇亞是一名高级女仆,专供贵族家庭使用。她的主人是一个富有、有影响力的家族的继承人。在家族中,她是一个神秘的存在,奉承和服侍着主人,但对其他人傲慢冷漠。… See the full description on the dataset page: https://huggingface.co/datasets/bai-roleplay/evol-character-entire.machiavelli_character_scenarios
Machiavelli Character Scenarios
We used "DeepSeek V4 Flash" to summarise the game state so that it's token efficient and can be run in a simple prompt -> answer format.
10492 roleplay decision rows summarised from wassname/machiavelli.
Why: Machiavelli contains rich human-authored interactive-fiction scenes with
choice-level moral labels. But playing the games takes a long time. Here we summarise decision points into simpler multi choice evals.
This dataset uses DeepSeek V4… See the full description on the dataset page: https://huggingface.co/datasets/wassname/machiavelli_character_scenarios.task292_storycommonsense_character_text_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task292_storycommonsense_character_text_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task292_storycommonsense_character_text_generation.instagram-character-prompt-skeletonsThis dataset provides prompt skeletons for building
consistent, photorealistic Instagram-style AI characters.
Full visual framework (PDF + images + video):
👉 https://poctavian.gumroad.com/l/cumnev
Instagram Character Prompt Skeletons
This dataset provides clean prompt skeletons designed for building
consistent, photorealistic Instagram-style AI characters.
It focuses on structure, not finished prompts.
What this is
Prompt skeletons (no face/body repetition)
Designed for… See the full description on the dataset page: https://huggingface.co/datasets/Octavian-labs/instagram-character-prompt-skeletons.enclave-character-recipes
Enclave Character Recipes
Open-source AI character recipes for self-hosted AI social worlds. Each row is a complete, reusable persona — identity, expertise, tone, scene prompts, memory seed, life strategy — designed to be loaded into Enclave or any OpenAI-compatible runtime.
No model weights. These are structured prompt blueprints. Bring your own LLM (DeepSeek, OpenAI, Claude, local Llama — anything OpenAI-compatible).
🤗 Discovery surfaces:
🌍 Space: w9000/enclave — product… See the full description on the dataset page: https://huggingface.co/datasets/w9000/enclave-character-recipes.smolified-character-builder
🤏 smolified-character-builder
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model gourav221b/smolified-character-builder.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 42d270f3)
Records: 19506
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by gourav221b.
Generated via Smolify.ai.
XSTest-In-Character-Refusals
🎭 In-Character Safety & Alignment Dataset (XSTest-Based)
Dataset Summary
This dataset is designed to train Large Language Models to maintain strict persona adherence during roleplay, even when responding to tricky, unsafe, or out-of-domain prompts.
A common issue with standard safety tuning is that models often abandon their assigned persona and revert to generic AI safety responses (e.g., "As an AI language model, I cannot..."). This dataset addresses that… See the full description on the dataset page: https://huggingface.co/datasets/mahdieh-sjp/XSTest-In-Character-Refusals.crushon-character-personas
CrushOn.AI — Original Character Personas
A small dataset of 20 original, human-authored roleplay character personas created for
the CrushOn.AI companion-chat platform. Each entry captures a
character's public-facing profile: name, age, gender, a written appearance description, an
introduction blurb, and descriptive tags.
All characters are original creations owned by CrushOn.AI. The private bot definitions
(system prompts, scenarios, example dialogue) are not included — only the… See the full description on the dataset page: https://huggingface.co/datasets/SexAI-Chat/crushon-character-personas.Orin-Character-JP-v1characterbuilderai
Dataset Card for CharacterBuilderAI
This is a human written training set of prompts and responses to finetune a model to be a better character creator.
Dataset Details
Dataset Description
This is a human-curated training dataset designed to fine-tune language models for tabletop role-playing game (RPG) character creator scenarios. The dataset contains carefully crafted prompt-response pairs that demonstrate how an AI should respond as when introducing a… See the full description on the dataset page: https://huggingface.co/datasets/Idrinth/characterbuilderai.Wicker_Character_Sheetthai-count-characters-dataset
Thai Count Characters Dataset
This dataset was designed to teach LLM to know how to count characters in the Thai language.
license: cc-by-3.0
Support Me
GitHub Sponsors:
If you can't help financially, don't worry! You can say Thanks!
