datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
polyu-storyworld-charactersMixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Mixamo-Animations-Characters.Mixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Mixamo-Animations-Characters.Kohaku-Delta-Alpha-Characters-Result
Test Index for Kohaku-Delta-Alpha Model
ID
Tag
Copyright
Gender
Posts
CCIP
AIC
BP
Core Tags
2075918
inkling_player_character
splatoon_(series)
female
9646
0.383333
0.998157
0.645165
inkling girl, long hair, tentacle hair, pointy ears, bangs, blunt bangs, red eyes
2040387
doodle_sensei_(blue_archive)
blue_archive
female
6632
0.721212
0.992072
0.625653
halo, bangs, blue eyes, breasts, long hair, black hair, blue hair, hair ornament
1978860
2b_(nier:automata)… See the full description on the dataset page: https://huggingface.co/datasets/AngelBottomless/Kohaku-Delta-Alpha-Characters-Result.game_characters
Database of Characters in Mobile Games
All the character in the following games are supported:
Arknights (crawled from https://prts.wiki)
Fate/Grand Order (crawled from https://fgo.wiki)
Azur Lane (crawled from https://wiki.biligame.com/blhx)
Girls' Front-Line (crawled from https://iopwiki.com/)
Genshin Impact (crawled from https://genshin-impact.fandom.com/ja/wiki/%E5%8E%9F%E7%A5%9E_Wiki)
The source code and python library is hosted on narugo1992/gchar, and the scheduled job is… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/game_characters.character_select_stand_alone_apphttps://github.com/mirabarukaso/character_select_stand_alone_app
Devanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/Devanagari-Characters-Image.mixamo-characters
Mixamo Characters Dataset
This dataset contains 3D character models from Mixamo in FBX format.
Contents
Format: FBX (Autodesk Filmbox)
Pose: T-Pose
Rig: Mixamo auto-rig (compatible with Mixamo animations)
Total characters: 108
File Structure
<character_name>.fbx
Usage
from huggingface_hub import hf_hub_download
# Download a specific character
filepath = hf_hub_download(
repo_id="your-username/mixamo-characters",
filename="Abe.fbx"… See the full description on the dataset page: https://huggingface.co/datasets/GbotHQ/mixamo-characters.Devanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/rhythmjain30/Devanagari-Characters-Image.generic_charactersfictional-characters-image-datasetHow to use
Here is how to use this dataset:
from datasets import load_dataset
dataset = load_dataset("gryffindor-ISWS/fictional-characters-image-dataset")
This repository contains fictional characters dataset constructed from Wikidata for the research project "Draw Me Like Your Triples: Leveraging Generative AI for the Completion of Wikidata". The project was conducted by Raia Abu Ahmad, Martin Critelli, Şefika Efeoğlu, Eleonora Mancini, Célian Ringwald and Xinyue Zhang under the… See the full description on the dataset page: https://huggingface.co/datasets/gryffindor-ISWS/fictional-characters-image-dataset.qwen-synth-characters-fused
qwen-synth-characters-fused
Every SFW row of AbstractPhil/qwen-synth-characters processed by the
qwen-test-runner 12-process fused extraction system: age gate (strict) → 3×caption
structuring (Qwen3.5-9B, slot-registry schema) → 12 deterministic vision task JSONs
(tasks_json) → FusedScene (fused_json: entities with mask-containment-owned
stratified attributes, relations with continuous offsets, counts, shared basin) →
deterministic fused prompt (prompt_fused).
Shards are… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters-fused.qwen-synth-characters
Qwen Synthetic Characters
A dataset of 60,847 fully synthetic (AI-generated) human portrait/character images produced
with Qwen-Image + the Qwen-Image-Lightning 4-step LoRA, with a prompt-augmentation policy
designed to give balanced demographics, diverse facial expressions, and varied attributes — and
to counter the base model's tendency to default to a narrow set of faces.
[!IMPORTANT]
These are not real people. Every image is generated by a diffusion model from a text… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters.modular_characterscharacter_similarity
character_similarity
This is a dataset used for training models to determine whether two anime images (containing only one person) depict the same character. The dataset includes the following versions:
Version
Filename
Characters
Images
Information
v0
images_v0.tar.xz
2059
162116
Crawled from zerochan.net, includes images of Arknights, Fate/Grand Order, Genshin Impact, Girls' Frontline, and Azur Lane, as well as over 1500 other game or anime characters. The images are… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/character_similarity.modular_characters_largefont_square_charactersmodular_charactersv2Roleplay-Anime-CharactersA small synthetic (mostly) SFW dataset of (mostly) one on one character RP. Focuses on anime and game characters etc. This dataset tries to leverage the information larger models know about the characters to play them in character better than would normally be possible with generic characters.
The situations are generally absurd so the model is forced to generalize. It focuses on teaching the model how to be proactive, creative, emotional and take existing characters it may know about and… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Roleplay-Anime-Characters.English_Characters_Images
English Characters Image Dataset (A-Z, a-z, 0-9)
This dataset contains high-resolution (128x128 pixels) grayscale images of English characters, including uppercase letters (A-Z), lowercase letters (a-z), and digits (0-9). Each character is available in 80,000 to 100,000 unique font styles, making it one of the most comprehensive resources for character-level image modeling.
Dataset Description
The images in this dataset have been generated by rendering over 85,000… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/English_Characters_Images.booru-characters
Booru Characters
Overview
A line-oriented JSON dataset of character tag metadata extracted from Danbooru using the Danbooru API. Each record contains tag-level metadata and simple relationships between tags.
Contents
hf_dataset/characters.jsonl: one JSON object per line. Each object contains the fields described below.
hf_dataset/dataset_info.json: minimal metadata describing the exported features.
Fields (per record)
id (int):… See the full description on the dataset page: https://huggingface.co/datasets/Sn0w123/booru-characters.synthetic-characters
Synthetic Characters Dataset
A synthetic image dataset generated with Flux Schnell featuring structured character prompts designed for training character generation, fashion understanding, and portrait synthesis models.
Recommended Filters:
Age
People Count
Hair color
Camera angle
Anime/Realistic
Nudity/Clothed
I'll likely recaption everything with a list of classifications attached to them for easy filtering later.
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-characters.Anime_Characters
📚 Dataset Summary
This dataset features 1,941 anime character images, neatly organized into 322 folders, each representing a different anime series 🎌.
📦 Size of downloaded files: 152 MB
🪄 Size of auto-converted Parquet files: 151 MB
📊 Split: Train only
🎭 Classes: 322 unique anime titles
Perfect for image classification, anime recommendation systems, and visual style analysis! 🎨✨
🏆 Supported Tasks
🖼️ Image Classification: Predict the anime title based on a… See the full description on the dataset page: https://huggingface.co/datasets/adi2606/Anime_Characters.Characterskuzushiji-dataset-characters
Dataset Card for Kuzushiji Character Dataset
Dataset Summary
The Kuzushiji Character Dataset contains individual character crops generated from full
page images and character-coordinate annotations. Crops retain their original pixel size;
they are not resized or padded. Each record keeps the source book, page, character ID,
Unicode label, original page box, and the clamped box actually used for cropping.
This card describes the generated repository… See the full description on the dataset page: https://huggingface.co/datasets/Kotomiya07/kuzushiji-dataset-characters.metrixel-rigged-characters
Metrixel Animated Character Samples
A small set of rigged, animated humanoid character models, free to use in your own projects. Each file ships with a skeleton, skin weights and a motion clip, so it drops straight into a scene — or retarget your own motion onto the rig instead.
They are published as complimentary sample content for Metrixel, a local 3D dataset-preparation toolchain — handy starting assets if you want to try a multi-view render, SDF or mesh-tensor export… See the full description on the dataset page: https://huggingface.co/datasets/EntVista/metrixel-rigged-characters.ai-characters-QA
データ作成方法
以下の質問集のデータセットを利用しました。
質問に対する回答をLLMによって生成し、QA集のデータセットを新たに作成しました。各種キャラクターは、回答生成時のシステムプロンプトでキャラ付けしています。
回答生成には、llama-cpp-pythonとunsloth/gemma-3-27b-it-UD-Q8_K_XL.ggufを使用しました。
aituber_question_dataset
Gemini-3.1-Pro-GLM5-CharactersPrompts generated by Gemini 3.1 Pro.
Responses generated by Gemini 3.1 Pro.
Reasoning traces:
Step 1: Generated by GLM 5 which was provided the original system prompt / knowledge
Step 2: Edited by Gemini to fix any contradictions with the existing response
Step 3: Edited again to remove / reduce drafting and remove reasoning related to safety / refusals
System prompt was generated based on the original and whatever constraints / rules etc. had been mentioned in the reasoning trace.
A small… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Gemini-3.1-Pro-GLM5-Characters.chub_popular_charactersDevanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/Yash141414/Devanagari-Characters-Image.
