datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
character_index
Anime Character Index
This dataset if for collecting all the hot characters from the internet, and extract their features and core tags. It will be useful for automatically testing the character generating ability of the anime-style base models.
7371 characters in total.
Copyrights
Copyright
Count
kantai_collection
393
pokemon
380
fate_(series)
350
hololive
277
blue_archive234
arknights
200
idolmaster
192
touhou
186
fire_emblem
168
umamusume… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/character_index.game_character_skins
Game Character Skins Dataset
Summary
This comprehensive dataset contains game character skins and artwork from multiple popular mobile and PC games, providing a rich collection of character visual assets for computer vision research and game development applications. The dataset spans eight major game titles including Arknights, Azur Lane, Blue Archive, Fate/Grand Order, Genshin Impact, Girls' Frontline, Neural Cloud, Nikke, Path to Nowhere, and Honkai: Star Rail… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/game_character_skins.common-crawl-character-countspolyu-storyworld-charactersMixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Mixamo-Animations-Characters.Mixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Mixamo-Animations-Characters.Kohaku-Delta-Alpha-Characters-Result
Test Index for Kohaku-Delta-Alpha Model
ID
Tag
Copyright
Gender
Posts
CCIP
AIC
BP
Core Tags
2075918
inkling_player_character
splatoon_(series)
female
9646
0.383333
0.998157
0.645165
inkling girl, long hair, tentacle hair, pointy ears, bangs, blunt bangs, red eyes
2040387
doodle_sensei_(blue_archive)
blue_archive
female
6632
0.721212
0.992072
0.625653
halo, bangs, blue eyes, breasts, long hair, black hair, blue hair, hair ornament
1978860
2b_(nier:automata)… See the full description on the dataset page: https://huggingface.co/datasets/AngelBottomless/Kohaku-Delta-Alpha-Characters-Result.game_characters
Database of Characters in Mobile Games
All the character in the following games are supported:
Arknights (crawled from https://prts.wiki)
Fate/Grand Order (crawled from https://fgo.wiki)
Azur Lane (crawled from https://wiki.biligame.com/blhx)
Girls' Front-Line (crawled from https://iopwiki.com/)
Genshin Impact (crawled from https://genshin-impact.fandom.com/ja/wiki/%E5%8E%9F%E7%A5%9E_Wiki)
The source code and python library is hosted on narugo1992/gchar, and the scheduled job is… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/game_characters.character_select_stand_alone_apphttps://github.com/mirabarukaso/character_select_stand_alone_app
Devanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/Devanagari-Characters-Image.generic_character_skins
Generic Character Skins Dataset
Summary
This comprehensive dataset provides an extensive collection of character images sourced from Zerochan across multiple popular anime, game, and manga franchises. The dataset contains meticulously organized character artwork spanning diverse genres including gacha games, idol franchises, fantasy series, and action RPGs. With over 2,000 character folders and thousands of high-quality images, this repository serves as a valuable… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/generic_character_skins.Hoyoverse_Character_Modelsmixamo-characters
Mixamo Characters Dataset
This dataset contains 3D character models from Mixamo in FBX format.
Contents
Format: FBX (Autodesk Filmbox)
Pose: T-Pose
Rig: Mixamo auto-rig (compatible with Mixamo animations)
Total characters: 108
File Structure
<character_name>.fbx
Usage
from huggingface_hub import hf_hub_download
# Download a specific character
filepath = hf_hub_download(
repo_id="your-username/mixamo-characters",
filename="Abe.fbx"… See the full description on the dataset page: https://huggingface.co/datasets/GbotHQ/mixamo-characters.Aesir-Character-CoT-roleplay
Overview
Think with your role.
Most reasoning datasets teach models to think like an AI. This one teaches them to think like the character.
Continue updating until money run out, I will try to update this dataset in near future
Stats
1,973 high-quality conversations (filtered from 2,000 distilled — 27 dropped: prohibited content + missing-review + empty-content)
~14,349 assistant turns, each with full character-POV reasoning
Teacher: deepseek-v4-pro… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/Aesir-Character-CoT-roleplay.characterDevanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/rhythmjain30/Devanagari-Characters-Image.comfyui-character-composer
ComfyUI Character Composer
Version: 2.0 • Release Date: 2026-05-18 • License: Apache-2.0
Structured, JSON-driven character prompt system for ComfyUI.
Designed for consistent, controllable character generation with Qwen-style workflows, replacing random prompt chaos with a more deterministic composition layer.
New / Advanced UI & Prompt generation (v2.0)
Character creation node now prints the final prompt it generates for easier debugging.
Character creation node… See the full description on the dataset page: https://huggingface.co/datasets/unh1nge/comfyui-character-composer.generic_charactersfictional-characters-image-datasetHow to use
Here is how to use this dataset:
from datasets import load_dataset
dataset = load_dataset("gryffindor-ISWS/fictional-characters-image-dataset")
This repository contains fictional characters dataset constructed from Wikidata for the research project "Draw Me Like Your Triples: Leveraging Generative AI for the Completion of Wikidata". The project was conducted by Raia Abu Ahmad, Martin Critelli, Şefika Efeoğlu, Eleonora Mancini, Célian Ringwald and Xinyue Zhang under the… See the full description on the dataset page: https://huggingface.co/datasets/gryffindor-ISWS/fictional-characters-image-dataset.moss-character-voices-bestof64
MOSS Character Voices — Best-of-64 (Stage 2)
Best-of-64 voice-acting takes from the 4.55B MOSS-TTS-Local voice-acting model
(laion/moss-tts-local-transformer-4.55b-voice-acting) for 13 evolved character voices.
Each prompt is a fixed, optimized champion performance direction (instruction) paired
with a Gemma-generated topic text (text) — together, one performance to render. For every
prompt we sample 64 takes with distinct seeds at 48 kHz, score each take, and rank the 64
within… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-bestof64.qwen-synth-characters-fused
qwen-synth-characters-fused
Every SFW row of AbstractPhil/qwen-synth-characters processed by the
qwen-test-runner 12-process fused extraction system: age gate (strict) → 3×caption
structuring (Qwen3.5-9B, slot-registry schema) → 12 deterministic vision task JSONs
(tasks_json) → FusedScene (fused_json: entities with mask-containment-owned
stratified attributes, relations with continuous offsets, counts, shared basin) →
deterministic fused prompt (prompt_fused).
Shards are… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters-fused.qwen-synth-characters
Qwen Synthetic Characters
A dataset of 60,847 fully synthetic (AI-generated) human portrait/character images produced
with Qwen-Image + the Qwen-Image-Lightning 4-step LoRA, with a prompt-augmentation policy
designed to give balanced demographics, diverse facial expressions, and varied attributes — and
to counter the base model's tendency to default to a narrow set of faces.
[!IMPORTANT]
These are not real people. Every image is generated by a diffusion model from a text… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters.mindbotz-character-sheets
MindBotz Character Reference Sheets
69 original AI-generated character reference sheets, built for the MindBotz animated character universe.
Each sheet contains turnaround views (front/side/back), 3 expression closeups, and annotation arrows with labels — proper animation-studio reference format.
Contents
images/ — 76 PNG files: 69 unique characters (+ a few bonus variants from re-renders)
characters.csv — index of character number → description
environments/ —… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/mindbotz-character-sheets.moss-character-voices-top3-captioned
MOSS Character Voices — Top-3 per Group, Captioned (training-ready)
The top-3 takes per group (rank 0/1/2 by reward model) of all ~12,900 groups in
laion/moss-character-voices-bestof64
— ~38,700 samples — each annotated with both LAION voice-acting caption pipelines and shipped with
all scores + metadata, ready to drop into training. Sharded as data/train-*-of-00129.parquet.
Captions (two pipelines, same clip)
caption_procedural — Procedural Voice Captions:
terse… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-top3-captioned.moss-character-reference-voices
MOSS character reference voices (1336 voices)
1336 distinct synthetic character voices, each mined from a cluster of generated MOSS-VA-v2 character
audio and auto-annotated by Gemini-3-Flash. For every cluster the model was shown the 3 cluster samples
their automatic voice scores, chose the single most representative sample, and wrote a full
casting-style profile.
Contents
dataset.jsonl — one row per voice: cid, name, tagline, description, age, gender, register… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-character-reference-voices.modular_charactersCharacterCodex
Dataset Card for Character Codex
Dataset Summary
The Character Codex is a comprehensive dataset featuring popular characters from a wide array of media types and genres. Each entry includes detailed information about the character, the media source, and a unique scenario involving the character. This dataset is valuable for synthetic data, RAG for generative AI, writers, game developers, and fans who want to explore and utilize rich character descriptions for various… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/CharacterCodex.Character_Loraskaidol-character-dataset
KAIdol Character Chat Dataset
한국어 캐릭터 롤플레이 대화 데이터셋
📋 목차
개요
데이터셋 통계
데이터 형식
캐릭터 목록
품질 지표
사용 방법
학습 가이드
제한사항
라이선스
🎯 개요
KAIdol Character Chat Dataset은 41개 고유 캐릭터의 롤플레이 대화 데이터셋입니다. 각 캐릭터는 독특한 **음성 프로필(Voice Profile)**을 가지고 있으며, 이를 기반으로 일관된 성격과 말투를 유지합니다.
주요 특징
특징
설명
🎭 41개 캐릭터
다양한 성격, 배경, 말투를 가진 캐릭터
🗣️ 음성 프로필
시그니처 표현, 종결어미, 금지 표현 정의
📊 3가지 형식
SFT, DPO, Multiturn 학습 지원
✅ 품질 검증
A등급 음성 프로필 일치율 (0.805)
🇰🇷 100% 한국어
자연스러운… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-character-dataset.character_similarity
character_similarity
This is a dataset used for training models to determine whether two anime images (containing only one person) depict the same character. The dataset includes the following versions:
Version
Filename
Characters
Images
Information
v0
images_v0.tar.xz
2059
162116
Crawled from zerochan.net, includes images of Arknights, Fate/Grand Order, Genshin Impact, Girls' Frontline, and Azur Lane, as well as over 1500 other game or anime characters. The images are… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/character_similarity.
