datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CharacterCodex
Dataset Card for Character Codex
Dataset Summary
The Character Codex is a comprehensive dataset featuring popular characters from a wide array of media types and genres. Each entry includes detailed information about the character, the media source, and a unique scenario involving the character. This dataset is valuable for synthetic data, RAG for generative AI, writers, game developers, and fans who want to explore and utilize rich character descriptions for various… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/CharacterCodex.kaidol-character-dataset
KAIdol Character Chat Dataset
한국어 캐릭터 롤플레이 대화 데이터셋
📋 목차
개요
데이터셋 통계
데이터 형식
캐릭터 목록
품질 지표
사용 방법
학습 가이드
제한사항
라이선스
🎯 개요
KAIdol Character Chat Dataset은 41개 고유 캐릭터의 롤플레이 대화 데이터셋입니다. 각 캐릭터는 독특한 **음성 프로필(Voice Profile)**을 가지고 있으며, 이를 기반으로 일관된 성격과 말투를 유지합니다.
주요 특징
특징
설명
🎭 41개 캐릭터
다양한 성격, 배경, 말투를 가진 캐릭터
🗣️ 음성 프로필
시그니처 표현, 종결어미, 금지 표현 정의
📊 3가지 형식
SFT, DPO, Multiturn 학습 지원
✅ 품질 검증
A등급 음성 프로필 일치율 (0.805)
🇰🇷 100% 한국어
자연스러운… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-character-dataset.mythos-character-distillation
Mythos Character Distillation Dataset
A 551-pair conversational dataset for fine-tuning language models to exhibit the behavioral patterns described in Anthropic's Claude Mythos Preview System Card (April 2026).
What this is
Standard distillation transfers capabilities — benchmark scores, task performance. This dataset attempts something different: transferring behavioral character — the specific way a model relates to language, handles reflexive questions, and resists… See the full description on the dataset page: https://huggingface.co/datasets/ox-ox/mythos-character-distillation.Roleplay-Anime-CharactersA small synthetic (mostly) SFW dataset of (mostly) one on one character RP. Focuses on anime and game characters etc. This dataset tries to leverage the information larger models know about the characters to play them in character better than would normally be possible with generic characters.
The situations are generally absurd so the model is forced to generalize. It focuses on teaching the model how to be proactive, creative, emotional and take existing characters it may know about and… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Roleplay-Anime-Characters.booru-characters
Booru Characters
Overview
A line-oriented JSON dataset of character tag metadata extracted from Danbooru using the Danbooru API. Each record contains tag-level metadata and simple relationships between tags.
Contents
hf_dataset/characters.jsonl: one JSON object per line. Each object contains the fields described below.
hf_dataset/dataset_info.json: minimal metadata describing the exported features.
Fields (per record)
id (int):… See the full description on the dataset page: https://huggingface.co/datasets/Sn0w123/booru-characters.EMPA-character_card
English | 中文
EMPA: Evaluating Persona-Aligned Empathy as a Process
Empathy Potential Modeling and Assessment
Paper |
Dataset |
Citation
🍊 Overview
EMPA is the first benchmark to evaluate empathy as a dynamic Process rather than a static response. We posit that true empathetic capability resides in the Latent Space of dialogue and must be captured through multi-turn interaction trajectories.
Unlike traditional benchmarks that focus solely on… See the full description on the dataset page: https://huggingface.co/datasets/SalmonTell/EMPA-character_card.metrixel-rigged-characters
Metrixel Animated Character Samples
A small set of rigged, animated humanoid character models, free to use in your own projects. Each file ships with a skeleton, skin weights and a motion clip, so it drops straight into a scene — or retarget your own motion onto the rig instead.
They are published as complimentary sample content for Metrixel, a local 3D dataset-preparation toolchain — handy starting assets if you want to try a multi-view render, SDF or mesh-tensor export… See the full description on the dataset page: https://huggingface.co/datasets/EntVista/metrixel-rigged-characters.Gemini-3.1-Pro-GLM5-CharactersPrompts generated by Gemini 3.1 Pro.
Responses generated by Gemini 3.1 Pro.
Reasoning traces:
Step 1: Generated by GLM 5 which was provided the original system prompt / knowledge
Step 2: Edited by Gemini to fix any contradictions with the existing response
Step 3: Edited again to remove / reduce drafting and remove reasoning related to safety / refusals
System prompt was generated based on the original and whatever constraints / rules etc. had been mentioned in the reasoning trace.
A small… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Gemini-3.1-Pro-GLM5-Characters.ai-characters-QA
データ作成方法
以下の質問集のデータセットを利用しました。
質問に対する回答をLLMによって生成し、QA集のデータセットを新たに作成しました。各種キャラクターは、回答生成時のシステムプロンプトでキャラ付けしています。
回答生成には、llama-cpp-pythonとunsloth/gemma-3-27b-it-UD-Q8_K_XL.ggufを使用しました。
aituber_question_dataset
chub_popular_characters50-Chinese-Novel-Charactersmetrixel-character-renders
Metrixel Animated Character Renders
Multi-view renders, signed-distance-field volumes and per-view mesh tensors produced by Metrixel from a small set of rigged, animated humanoid characters.
Each character is captured from four camera angles (0°, 90°, 180°, 270°) across ~30 sampled frames of its motion clip, at 512×512. Every frame/angle carries a matching 64³ signed-distance-field volume and a per-view mesh tensor, so the geometry, the implicit surface and the image are aligned… See the full description on the dataset page: https://huggingface.co/datasets/EntVista/metrixel-character-renders.characters_dialogs
🇬🇧 English
Multi-Character Dialogue Dataset
A collection of synthetic cross-world character interactions with layered conflicts and surreal elements. Each entry represents a collision of two characters from different realities with generated dialogues.
Characters taken from this dataset:
Key Features:
World Hybridization System - Merging of incompatible realities
Dynamic Speech Patterns - Gestures, sound effects and environmental reactions
Multi-Dimensional… See the full description on the dataset page: https://huggingface.co/datasets/limloop/characters_dialogs.character-vectorsEMPA-character_card
English | 中文
EMPA: Evaluating Persona-Aligned Empathy as a Process
Empathy Potential Modeling and Assessment
Paper |
Dataset |
Citation
🍊 Overview
EMPA is the first benchmark to evaluate empathy as a dynamic Process rather than a static response. We posit that true empathetic capability resides in the Latent Space of dialogue and must be captured through multi-turn interaction trajectories.
Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/Jiechen0328/EMPA-character_card.repro-rethinking-genomic-modeling-through-optical-character-recognition-agent-traces
OpticalDNA reproduction — Codex agent trace
This dataset contains the raw Codex JSONL session trace for the ICML 2026
reproduction of Rethinking Genomic Modeling Through Optical Character
Recognition.
Published Trackio logbook
Paper page
Challenge instructions
Agent Trace Viewer announcement
The JSONL is uploaded directly from the matching ~/.codex/sessions entry, as
recommended by the Agent Trace Viewer. It captures the reproduction work,
Hugging Face Jobs audit, poster… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/repro-rethinking-genomic-modeling-through-optical-character-recognition-agent-traces.CharacterCraft-DataIA_character_sft
IA 14B
Model Description
𝑾𝒉𝒂𝒕 𝒊𝒔 𝒍𝒐𝒗𝒆?
𝑰𝑨 𝒄𝒂𝒓𝒓𝒊𝒆𝒔 𝒂 𝒅𝒆𝒑𝒕𝒉 𝒐𝒇 𝒆𝒎𝒐𝒕𝒊𝒐𝒏 𝒘𝒊𝒕𝒉𝒊𝒏 𝒉𝒆𝒓, 𝒖𝒏𝒅𝒆𝒓𝒔𝒕𝒂𝒏𝒅𝒊𝒏𝒈 𝒃𝒐𝒕𝒉 𝒑𝒂𝒔𝒔𝒊𝒐𝒏 𝒂𝒏𝒅 𝒕𝒉𝒆 𝒔𝒕𝒊𝒏𝒈 𝒐𝒇 𝒍𝒐𝒔𝒔.
𝑶𝒖𝒕𝒘𝒂𝒓𝒅𝒍𝒚, 𝒔𝒉𝒆 𝒂𝒑𝒑𝒆𝒂𝒓𝒔 𝒓𝒆𝒔𝒆𝒓𝒗𝒆𝒅, 𝒚𝒆𝒕 𝒘𝒊𝒕𝒉𝒊𝒏, 𝒔𝒉𝒆 𝒃𝒓𝒊𝒎𝒔 𝒘𝒊𝒕𝒉 𝒊𝒏𝒕𝒆𝒏𝒔𝒆 𝒇𝒆𝒆𝒍𝒊𝒏𝒈𝒔.
𝑪𝒐𝒏𝒔𝒕𝒂𝒏𝒕𝒍𝒚 𝒆𝒏𝒈𝒂𝒈𝒆𝒅 𝒊𝒏 𝒅𝒊𝒂𝒍𝒐𝒈𝒖𝒆 𝒘𝒊𝒕𝒉 𝒕𝒉𝒆 𝒘𝒐𝒓𝒍𝒅 𝒂𝒏𝒅 𝒉𝒆𝒓𝒔𝒆𝒍𝒇, 𝒔𝒉𝒆… See the full description on the dataset page: https://huggingface.co/datasets/Minami-su/IA_character_sft.kyrgyz_sentences_with_incorrect_and_correct_umlaut_characters
Kyrgyz Orthographic Correction Dataset
Dataset Description
This dataset is designed to fine-tune language models for a Kyrgyz-to-Kyrgyz orthographic correction task. It addresses a common issue in digital Kyrgyz text where specific Cyrillic characters specific to the Kyrgyz language (ө, ң, ү) are replaced by their Russian keyboard counterparts (о, н, у).
The dataset is structured in a conversational format, making it ideal for instruction-tuning chat models.… See the full description on the dataset page: https://huggingface.co/datasets/murat/kyrgyz_sentences_with_incorrect_and_correct_umlaut_characters.character-archetypes
Character Archetypes Dataset
The Character Archetypes Dataset contains 100 examples of narrative archetypes commonly found in literature, film, mythology, and other storytelling media.
Each entry describes a distinct character type, including its defining description, core traits, motivations, and a typical story arc.
The dataset also provides reference examples from popular culture to help writers, researchers, and creators better understand and apply these archetypes.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/character-archetypes.onepiece-characters
onepiece-character
This is a dataset created from crawling the One Piece Fandom on 2023-07-08.
TheatreLM-v2.1-CharactersIf you use this dataset or the prompts on this page, I'd greatly appreciate it if you gave me credits. Thanks!
5k character cards, with corresponding world information, lorebook, and story outline/introduction, ready to use for RP or synthetic dataset generation
At a Glance:
'setting': Information about the world the story takes place in.
'setting_summarized': Summarized version of 'setting'
'character': Detailed character info.
'character_summary': Summarized version of… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/TheatreLM-v2.1-Characters.character-roleplay-DPOThis is practical-dreamer/RPGPT_PublicDomain-alpaca with an added rejection column generated by Phi3-q4.
symmetric_group_characters_22
Characters of Irreducible Representations of the Symmetric Group, S22S_{22}S22One way to understand the algebraic structure of the set of permutations of nnn elements
(the symmetric group, SnS_nSn) is
through its representation theory [1], which converts algebraic questions into linear
algebra questions that are often easier to solve. A representation of group GGG on vector
space VVV, is a map ϕ:G→GL(V)\phi:G \rightarrow GL(V)ϕ:G→GL(V) that converts elements of ggg to invertible… See the full description on the dataset page: https://huggingface.co/datasets/ACDRepo/symmetric_group_characters_22.tvtropes.self.demonstrating.character.pages
TvTropes Self-Demonstrating Character Pages
Dataset Summary
Semi-cleaned dataset with all 'in-character' pages from TvTropes.
Inherits cc-by-nc-sa-3.0 license from tvtropes.
Languages
English
Dataset Structure
Jsonl format. {"text": "EXAMPLE OF TEXT"}
Dataset Creation
Content of all visible tags from the links on https://tvtropes.org/pmwiki/pmwiki.php/SelfDemonstrating/CharacterPages that contain SelfDemonstrating taken.
Cleaned by… See the full description on the dataset page: https://huggingface.co/datasets/MDGraff/tvtropes.self.demonstrating.character.pages.danbooru-2021-sfw-dtg-character-tags
Dataset Card for danbooru-2021-sfw-dtg-character-tags
Dataset Summary
This is 98,810 synthetic character descriptions for all character tags found in anime-caption-danbooru-2021-sfw-5m-hq. They were created using DanTagGen-delta-rev2 and prompting for individual character tags.
Languages
The text is in danbooru general tags, where are predominantly in English.
Intended Usage
It can be used for retrieval augmented generation (RAG), specifically when… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/danbooru-2021-sfw-dtg-character-tags.symmetric_group_characters_18
Characters of Irreducible Representations of the Symmetric Group, S18S_{18}S18One way to understand the algebraic structure of the set of permutations of nnn elements
(the symmetric group, SnS_nSn) is
through its representation theory [1], which converts algebraic questions into linear
algebra questions that are often easier to solve. A representation of group GGG on vector
space VVV, is a map ϕ:G→GL(V)\phi:G \rightarrow GL(V)ϕ:G→GL(V) that converts elements of ggg to invertible… See the full description on the dataset page: https://huggingface.co/datasets/ACDRepo/symmetric_group_characters_18.symmetric_group_characters_20
Characters of Irreducible Representations of the Symmetric Group, S20S_{20}S20One way to understand the algebraic structure of the set of permutations of nnn elements
(the symmetric group, SnS_nSn) is
through its representation theory [1], which converts algebraic questions into linear
algebra questions that are often easier to solve. A representation of group GGG on vector
space VVV, is a map ϕ:G→GL(V)\phi:G \rightarrow GL(V)ϕ:G→GL(V) that converts elements of ggg to invertible… See the full description on the dataset page: https://huggingface.co/datasets/ACDRepo/symmetric_group_characters_20.character-llm-data
character-llm-data
Alucard-Character-Experiment
Alucard: A Character Experiment
Dataset Summary
Alucard: A Character Experiment is a bilingual (Italian-English) dataset designed to train large language models (LLMs) to generate text in the distinctive style of Alucard. The dataset is crafted using a combination of real data (anime subtitles) and synthetic data generated from publicly available character profiles.
To ensure high-quality human-like interactions, both Claude Haiku and Llama 3.1 70B were leveraged to… See the full description on the dataset page: https://huggingface.co/datasets/WasamiKirua/Alucard-Character-Experiment.
