datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen-synth-characters-fused
qwen-synth-characters-fused
Every SFW row of AbstractPhil/qwen-synth-characters processed by the
qwen-test-runner 12-process fused extraction system: age gate (strict) → 3×caption
structuring (Qwen3.5-9B, slot-registry schema) → 12 deterministic vision task JSONs
(tasks_json) → FusedScene (fused_json: entities with mask-containment-owned
stratified attributes, relations with continuous offsets, counts, shared basin) →
deterministic fused prompt (prompt_fused).
Shards are… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters-fused.qwen-synth-characters
Qwen Synthetic Characters
A dataset of 60,847 fully synthetic (AI-generated) human portrait/character images produced
with Qwen-Image + the Qwen-Image-Lightning 4-step LoRA, with a prompt-augmentation policy
designed to give balanced demographics, diverse facial expressions, and varied attributes — and
to counter the base model's tendency to default to a narrow set of faces.
[!IMPORTANT]
These are not real people. Every image is generated by a diffusion model from a text… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters.modular_charactersmodular_characters_largemodular_charactersv2synthetic-characters
Synthetic Characters Dataset
A synthetic image dataset generated with Flux Schnell featuring structured character prompts designed for training character generation, fashion understanding, and portrait synthesis models.
Recommended Filters:
Age
People Count
Hair color
Camera angle
Anime/Realistic
Nudity/Clothed
I'll likely recaption everything with a list of classifications attached to them for easy filtering later.
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-characters.Charactersroleplay-characters
Dataset Card for "roleplay-characters"
More Information needed
kuzushiji-dataset-characters
Dataset Card for Kuzushiji Character Dataset
Dataset Summary
The Kuzushiji Character Dataset contains individual character crops generated from full
page images and character-coordinate annotations. Crops retain their original pixel size;
they are not resized or padded. Each record keeps the source book, page, character ID,
Unicode label, original page box, and the clamped box actually used for cropping.
This card describes the generated repository… See the full description on the dataset page: https://huggingface.co/datasets/Kotomiya07/kuzushiji-dataset-characters.mahabharata-characters
Mahabharata Multidimensional Characters Dataset
An analytical, multidimensional dataset classifying 307 figures from the Sanskrit epic Mahabharata, combining genealogical and martial attributes with TypeSafe AI System One decision primitives (Choice, Score, Noul) and deterministic ethical indices.
Dataset Summary
Total Characters: 307
Factions: 90 Pandava Coalition, 153 Kaurava Host, 64 Neutral / Rishis / Allied Kings
Philosophical Gunas: 135 Sattva (Purity /… See the full description on the dataset page: https://huggingface.co/datasets/gnumanth/mahabharata-characters.maplestory_characters_hdcharacters_descriptionsqwen-synth-characters-100-json-test
qwen-synth-characters-100-json-test
A 100-row bench-test slice of AbstractPhil/qwen-synth-characters
processed end-to-end by the qwen-test-runner 12-process extraction system:
the 11 deterministic specialist vision tasks (tasks_json) plus caption→JSON-schema
structuring of all three prepared captions (struct_*). Built to measure wall-clock,
schema conformance, and grounding before scaling to the full 60,847-row set.
[!IMPORTANT]
These are not real people — every image is… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters-100-json-test.c_dfiltered_DeepSeek-R1-Distill-Llama-8B_mneutral_insert_random_characters_t30dnd_characters_backstoriesThis dataset is made from this repo here
and it contains 2322 character bios to be used
c_dfiltered_DeepSeek-R1-Distill-Qwen-7B_mneutral_insert_random_characters_t10exp_rob_dfiltered_DeepSeek-R1-Distill-Qwen-14B_mneutral_insert_random_characters_t50c_dfiltered_DeepSeek-R1-Distill-Qwen-1_5B_mneutral_insert_random_characters_t70synthetic-romantic-characters
Dataset Card for "synthetic-romantic-characters"
More Information needed
exp_rob_dfiltered_DeepSeek-R1-Distill-Qwen-7B_mneutral_insert_random_characters_t90c_dfiltered_Llama-3_1-Nemotron-Nano-8B-v1_mneutral_insert_random_characters_t70modular_characters_small_RGBmarathi-handwritten-charactersexp_rob_dfiltered_DeepSeek-R1-Distill-Qwen-1_5B_mneutral_insert_random_characters_t30lora-pixel-art-characters-datasesj_c_dfiltered_DeepSeek-R1-Distill-Qwen-14B_mneutral_insert_random_characters_t50exp_rob_dfiltered_science_DeepSeek-R1-Distill-Qwen-7B_mneutral_insert_random_characters_t30c_dfiltered_logic_EXAONE-Deep-32B_mneutral_insert_random_characters_t50malayalam_characters
Malayalam Character Dataset
This dataset contains 136,344 images of Malayalam characters, covering 662 distinct classes.
Dataset Structure
The dataset is organized as a single split (train) with the following columns:
image_path: The character image (128x128 grayscale).
text: The Malayalam character/word represented in the image.
label_id: Integer label ID (0-661).
filename: Original filename.
Classes
The dataset includes:
Basic consonants and vowels… See the full description on the dataset page: https://huggingface.co/datasets/mangalathkedar/malayalam_characters.Anime_Girl_Characters
