datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
polyu-storyworld-charactersMixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Mixamo-Animations-Characters.Mixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Mixamo-Animations-Characters.qwen-synth-characters-fused
qwen-synth-characters-fused
Every SFW row of AbstractPhil/qwen-synth-characters processed by the
qwen-test-runner 12-process fused extraction system: age gate (strict) → 3×caption
structuring (Qwen3.5-9B, slot-registry schema) → 12 deterministic vision task JSONs
(tasks_json) → FusedScene (fused_json: entities with mask-containment-owned
stratified attributes, relations with continuous offsets, counts, shared basin) →
deterministic fused prompt (prompt_fused).
Shards are… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters-fused.qwen-synth-characters
Qwen Synthetic Characters
A dataset of 60,847 fully synthetic (AI-generated) human portrait/character images produced
with Qwen-Image + the Qwen-Image-Lightning 4-step LoRA, with a prompt-augmentation policy
designed to give balanced demographics, diverse facial expressions, and varied attributes — and
to counter the base model's tendency to default to a narrow set of faces.
[!IMPORTANT]
These are not real people. Every image is generated by a diffusion model from a text… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters.modular_characterscharacter_similarity
character_similarity
This is a dataset used for training models to determine whether two anime images (containing only one person) depict the same character. The dataset includes the following versions:
Version
Filename
Characters
Images
Information
v0
images_v0.tar.xz
2059
162116
Crawled from zerochan.net, includes images of Arknights, Fate/Grand Order, Genshin Impact, Girls' Frontline, and Azur Lane, as well as over 1500 other game or anime characters. The images are… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/character_similarity.modular_characters_largemodular_charactersv2font_square_charactersRoleplay-Anime-CharactersA small synthetic (mostly) SFW dataset of (mostly) one on one character RP. Focuses on anime and game characters etc. This dataset tries to leverage the information larger models know about the characters to play them in character better than would normally be possible with generic characters.
The situations are generally absurd so the model is forced to generalize. It focuses on teaching the model how to be proactive, creative, emotional and take existing characters it may know about and… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Roleplay-Anime-Characters.booru-characters
Booru Characters
Overview
A line-oriented JSON dataset of character tag metadata extracted from Danbooru using the Danbooru API. Each record contains tag-level metadata and simple relationships between tags.
Contents
hf_dataset/characters.jsonl: one JSON object per line. Each object contains the fields described below.
hf_dataset/dataset_info.json: minimal metadata describing the exported features.
Fields (per record)
id (int):… See the full description on the dataset page: https://huggingface.co/datasets/Sn0w123/booru-characters.synthetic-characters
Synthetic Characters Dataset
A synthetic image dataset generated with Flux Schnell featuring structured character prompts designed for training character generation, fashion understanding, and portrait synthesis models.
Recommended Filters:
Age
People Count
Hair color
Camera angle
Anime/Realistic
Nudity/Clothed
I'll likely recaption everything with a list of classifications attached to them for easy filtering later.
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-characters.Characterskuzushiji-dataset-characters
Dataset Card for Kuzushiji Character Dataset
Dataset Summary
The Kuzushiji Character Dataset contains individual character crops generated from full
page images and character-coordinate annotations. Crops retain their original pixel size;
they are not resized or padded. Each record keeps the source book, page, character ID,
Unicode label, original page box, and the clamped box actually used for cropping.
This card describes the generated repository… See the full description on the dataset page: https://huggingface.co/datasets/Kotomiya07/kuzushiji-dataset-characters.metrixel-rigged-characters
Metrixel Animated Character Samples
A small set of rigged, animated humanoid character models, free to use in your own projects. Each file ships with a skeleton, skin weights and a motion clip, so it drops straight into a scene — or retarget your own motion onto the rig instead.
They are published as complimentary sample content for Metrixel, a local 3D dataset-preparation toolchain — handy starting assets if you want to try a multi-view render, SDF or mesh-tensor export… See the full description on the dataset page: https://huggingface.co/datasets/EntVista/metrixel-rigged-characters.Gemini-3.1-Pro-GLM5-CharactersPrompts generated by Gemini 3.1 Pro.
Responses generated by Gemini 3.1 Pro.
Reasoning traces:
Step 1: Generated by GLM 5 which was provided the original system prompt / knowledge
Step 2: Edited by Gemini to fix any contradictions with the existing response
Step 3: Edited again to remove / reduce drafting and remove reasoning related to safety / refusals
System prompt was generated based on the original and whatever constraints / rules etc. had been mentioned in the reasoning trace.
A small… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Gemini-3.1-Pro-GLM5-Characters.ai-characters-QA
データ作成方法
以下の質問集のデータセットを利用しました。
質問に対する回答をLLMによって生成し、QA集のデータセットを新たに作成しました。各種キャラクターは、回答生成時のシステムプロンプトでキャラ付けしています。
回答生成には、llama-cpp-pythonとunsloth/gemma-3-27b-it-UD-Q8_K_XL.ggufを使用しました。
aituber_question_dataset
chub_popular_charactersroleplay-characters
Dataset Card for "roleplay-characters"
More Information needed
50-Chinese-Novel-Characterscharacters_dialogs
🇬🇧 English
Multi-Character Dialogue Dataset
A collection of synthetic cross-world character interactions with layered conflicts and surreal elements. Each entry represents a collision of two characters from different realities with generated dialogues.
Characters taken from this dataset:
Key Features:
World Hybridization System - Merging of incompatible realities
Dynamic Speech Patterns - Gestures, sound effects and environmental reactions
Multi-Dimensional… See the full description on the dataset page: https://huggingface.co/datasets/limloop/characters_dialogs.maplestory_characters_hdanime-characters-datasetkyrgyz_sentences_with_incorrect_and_correct_umlaut_characters
Kyrgyz Orthographic Correction Dataset
Dataset Description
This dataset is designed to fine-tune language models for a Kyrgyz-to-Kyrgyz orthographic correction task. It addresses a common issue in digital Kyrgyz text where specific Cyrillic characters specific to the Kyrgyz language (ө, ң, ү) are replaced by their Russian keyboard counterparts (о, н, у).
The dataset is structured in a conversational format, making it ideal for instruction-tuning chat models.… See the full description on the dataset page: https://huggingface.co/datasets/murat/kyrgyz_sentences_with_incorrect_and_correct_umlaut_characters.Handwritten-characterspulpfiction-characters-photosmahabharata-characters
Mahabharata Multidimensional Characters Dataset
An analytical, multidimensional dataset classifying 307 figures from the Sanskrit epic Mahabharata, combining genealogical and martial attributes with TypeSafe AI System One decision primitives (Choice, Score, Noul) and deterministic ethical indices.
Dataset Summary
Total Characters: 307
Factions: 90 Pandava Coalition, 153 Kaurava Host, 64 Neutral / Rishis / Allied Kings
Philosophical Gunas: 135 Sattva (Purity /… See the full description on the dataset page: https://huggingface.co/datasets/gnumanth/mahabharata-characters.allaimovies-ai-characters
allaimovies AI characters
3,263 artificial-intelligence characters from 1,884 science-fiction films (1911-2026):
robots, androids, cyborg intelligences, sentient computers, virtual humans and uploaded minds, each
with its kind, how the film presents its gender, whether it helps or opposes the humans, and how
prominent it is. Companion to the allaimovies film dataset from
https://github.com/prateek-0-gupta/allaimovies.
How it was made
For each film in the analysis… See the full description on the dataset page: https://huggingface.co/datasets/prateek-0-gupta/allaimovies-ai-characters.c_dfiltered_DeepSeek-R1-Distill-Qwen-7B_mneutral_insert_random_characters_t10
