datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aesir-Character-CoT-roleplay
Overview
Think with your role.
Most reasoning datasets teach models to think like an AI. This one teaches them to think like the character.
Continue updating until money run out, I will try to update this dataset in near future
Stats
1,973 high-quality conversations (filtered from 2,000 distilled — 27 dropped: prohibited content + missing-review + empty-content)
~14,349 assistant turns, each with full character-POV reasoning
Teacher: deepseek-v4-pro… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/Aesir-Character-CoT-roleplay.roleplay-bench
RP-Bench: Roleplay Quality Benchmark for LLMs
A multi-dimensional evaluation framework for measuring how well LLMs perform in roleplay scenarios — not just writing quality, but character consistency, user agency respect, lorebook integration, temporal reasoning, and genre-specific craft.
The LLM-as-judge signals in this benchmark disagree with real users about half the time. We're calibrating against human preferences via a public blind-arena. Help out at arena.l3vi4th4n.ai — each… See the full description on the dataset page: https://huggingface.co/datasets/lazyweasel/roleplay-bench.role-play-bench
Role-play Benchmark
A comprehensive benchmark for evaluating Role-play Agents in Chinese and English scenarios.
Dataset Summary
Role-play Benchmark is designed to evaluate Role-play Agents' ability to deliver immersive role-play experiences through Situated Reenactment. Unlike traditional benchmarks with verifiable answers, Role-play is fundamentally non-verifiable, e.g., there's no single "correct" response when a tsundere character is asked "Do you like me?".… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/role-play-bench.zhtw-roleplay-space-grimoire
Space Grimoire RP Corpus (Traditional Chinese)
Speaker-attributed dialogue from the original novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage; 283 chapters, ~1.7M characters), cut into scenes and assembled into ShareGPT-style role-play training data. The novel and this dataset are the work of 睡半夜怎麼三更, who holds the copyright and has no exclusive platform agreement. Data: CC BY 4.0. Code: Apache 2.0.
中文說明在下方
Dataset Summary
Source text
283… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/zhtw-roleplay-space-grimoire.role-play-bench
Role-play Benchmark
A comprehensive benchmark for evaluating Role-play Agents in Chinese and English scenarios.
Dataset Summary
Role-play Benchmark is designed to evaluate Role-play Agents' ability to deliver immersive role-play experiences through Situated Reenactment. Unlike traditional benchmarks with verifiable answers, Role-play is fundamentally non-verifiable, e.g., there's no single "correct" response when a tsundere character is asked "Do you like me?". Instead… See the full description on the dataset page: https://huggingface.co/datasets/EnlistedGhost/role-play-bench.Moral-RolePlay
Moral RolePlay
Paper | Code & Project Page
Abstract
Large Language Models (LLMs) are increasingly tasked with creative generation, including the simulation of fictional characters. However, their ability to portray non-prosocial, antagonistic personas remains largely unexamined. We hypothesize that the safety alignment of modern LLMs creates a fundamental conflict with the task of authentically role-playing morally ambiguous or villainous characters. To investigate this… See the full description on the dataset page: https://huggingface.co/datasets/Zihao1/Moral-RolePlay.RolePlay_Collection_random_ShareGPT
RolePlay_Collection_random_ShareGPT
This collection of random datasets for roleplay has been gathered from various sources. It has been processed, cleaned, and grammar-checked. However, it still requires significant work to be fully usable. Additional cleaning is necessary. I hope this proves helpful to someone.
gpt-roleplay-realm-Embeddings
GPT Roleplay Realm Embeddings
Embeddings of IlyaGusev/gpt_roleplay_realm, produced with amkdg/Qwen3-Embedding-8B-NVFP4 — 4096-d,
L2-normalized float16 (cosine = dot product).
8,700 conversations → 8,700 vectors
emb.npy — float16 [8700, 4096]
meta.parquet — one row per vector, aligned with emb.npy: id, uuid, tag, chunk, n_chunks, count, source_ref
manifest.json — counts and provenance
Usage
import numpy as np, pyarrow.parquet as pq
emb = np.load("emb.npy"… See the full description on the dataset page: https://huggingface.co/datasets/amkdg/gpt-roleplay-realm-Embeddings.chinese-roleplayroleplay_dialogues_extracted
Dialogues have not been extracted yet!
But characters do have been.
factual-multiagent-roleplay-ft-ru
march228/factual-multiagent-roleplay-ft-ru
Небольшой русскоязычный synthetic finetuning dataset для обучения модели следованию ролевым системным инструкциям личности при сохранении фактической опоры на контекст.
Что это за датасет
Этот набор сделан как instruction / finetuning dataset, а не как benchmark.
В каждой записи есть:
плотный system с персоной и тоном;
context, на который нужно опираться;
пользовательский question;
внутренние thoughts;
финальный answer.… See the full description on the dataset page: https://huggingface.co/datasets/march228/factual-multiagent-roleplay-ft-ru.chai-roleplay-labelednsfw: true means safe, false means nsfw
context_ratio: content between ** / total char
total_token_length: total amount of token from memory,prompt, chat_history, simply added, didn't remove repetitive tokens
GPTeacher_roleplay_standardized
Dataset Card for "GPTeacher_roleplay_standardized"
More Information needed
pippa_roleplay_standardizedairoboros_stage_3_roleplay_none_response_gpt-4o-inst_gpt_4o-mini_respGPTeacher_roleplay_standardized_cluster_0
Dataset Card for "GPTeacher_roleplay_standardized_cluster_0"
More Information needed
GPTeacher_roleplay_standardized_cluster_2_std
Dataset Card for "GPTeacher_roleplay_standardized_cluster_2_std"
More Information needed
GPTeacher_roleplay_standardized_cluster_1
Dataset Card for "GPTeacher_roleplay_standardized_cluster_1"
More Information needed
vicgalle__Roleplay-Llama-3-8Brole-play-bench
Role-play Benchmark
A comprehensive benchmark for evaluating Role-play Agents in Chinese and English scenarios.
Dataset Summary
Role-play Benchmark is designed to evaluate Role-play Agents' ability to deliver immersive role-play experiences through Situated Reenactment. Unlike traditional benchmarks with verifiable answers, Role-play is fundamentally non-verifiable, e.g., there's no single "correct" response when a tsundere character is asked "Do you like me?". Instead… See the full description on the dataset page: https://huggingface.co/datasets/web3w/role-play-bench.ELMB-RolePlayroleplay_dialogue_test
TX Roleplay Benchmark (Test)
This is a test dataset with sample data for demonstration.
Evaluation Dimensions
Worlds Dimensions
Basics: Text quality — avoiding garbled text, language mixing, severe repetition
Logic: Logical consistency — preventing character confusion, spatial errors, pronoun reference confusion
Knowledge: Factual reliability — balancing rules between virtual and real worlds
Stories Dimensions
Diversity: Expression variety and… See the full description on the dataset page: https://huggingface.co/datasets/ShanqiL/roleplay_dialogue_test.GPTeacher_roleplay_standardized_cluster_2
Dataset Card for "GPTeacher_roleplay_standardized_cluster_2"
More Information needed
vicgalle__Roleplay-Llama-3-8B-details
Dataset Card for Evaluation run of vicgalle/Roleplay-Llama-3-8B
Dataset automatically created during the evaluation run of model vicgalle/Roleplay-Llama-3-8B
The dataset is composed of 43 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vicgalle__Roleplay-Llama-3-8B-details.bunnycore__SmolLM2-1.7B-roleplay-lora-details
Dataset Card for Evaluation run of bunnycore/SmolLM2-1.7B-roleplay-lora
Dataset automatically created during the evaluation run of model bunnycore/SmolLM2-1.7B-roleplay-lora
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__SmolLM2-1.7B-roleplay-lora-details.fhai50032__RolePlayLake-7B-details
Dataset Card for Evaluation run of fhai50032/RolePlayLake-7B
Dataset automatically created during the evaluation run of model fhai50032/RolePlayLake-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fhai50032__RolePlayLake-7B-details.GPTeacher_roleplay_standardized_cluster_0_std
Dataset Card for "GPTeacher_roleplay_standardized_cluster_0_std"
More Information needed
GPTeacher_roleplay_standardized_cluster_1_std
Dataset Card for "GPTeacher_roleplay_standardized_cluster_1_std"
More Information needed
roleplay-test
roleplay-test
Test split for lm-eval-harness.
airoboros_stage_2_roleplay_only_none_resp_gpt-4o_gte-large
