datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
comfyui-character-composer
ComfyUI Character Composer
Version: 2.0 • Release Date: 2026-05-18 • License: Apache-2.0
Structured, JSON-driven character prompt system for ComfyUI.
Designed for consistent, controllable character generation with Qwen-style workflows, replacing random prompt chaos with a more deterministic composition layer.
New / Advanced UI & Prompt generation (v2.0)
Character creation node now prints the final prompt it generates for easier debugging.
Character creation node… See the full description on the dataset page: https://huggingface.co/datasets/unh1nge/comfyui-character-composer.brief-composer-sft-v1
BriefComposer SFT
Multi-image analytical brief rows composed from completed FireWatch, OceanScout, LandShift, and FloodPulse dataset folders (metadata/ + images/). Each sample stitches 1–4 images and metadata-derived headlines into one executive-style assistant reply.
Record counts (this build)
Split
JSONL lines
train
6307
validation
851
test
842
total
8000
Inputs
Source roots: one or more --source-root directories (each must contain… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/brief-composer-sft-v1.heb-composed-3style-hh104-hh1-hh2heb-connected-composed-4styleheb-composed-4style-hh104-hh1-hh2-hh3heb-bigram-composed-1mheb-connected-composed-3trackgeometric_spatial_compose_reasoningheb-composed-5style-hh104-hh1-hh2-hh3-hh105-BYcomfyui-character-composer
AIO Qwen Workflow
The repository now includes:
AIO Comfyui-Character-Composer Qwen Workflow.json
A unified all-in-one Qwen workflow designed around Character Composer.
This is no longer just:
a custom node
a tag helper
a prompt formatter
It is effectively a lightweight procedural character-generation system for ComfyUI.
The AIO workflow combines:
Qwen image editing
text-to-image
image-to-image
structured prompting
character consistency tools
composition preservation… See the full description on the dataset page: https://huggingface.co/datasets/BanjoBumpkin/comfyui-character-composer.brief-composer-sft-v1-duplicate
BriefComposer SFT
Multi-image analytical brief rows composed from completed FireWatch, OceanScout, LandShift, and FloodPulse dataset folders (metadata/ + images/). Each sample stitches 1–4 images and metadata-derived headlines into one executive-style assistant reply.
Record counts (this build)
Split
JSONL lines
train
2355
validation
310
test
335
total
3000
Inputs
Source roots: one or more --source-root directories (each must contain… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/brief-composer-sft-v1-duplicate.heb-composed-6style-hh104-hh1-hh2-hh3-hh105-hh106-BYCompoSET
CompoSET: A Controlled Single-Edit Benchmark for Vision-Language Compositionality
CompoSET contains 1,776 instances across 80 naturalistic scenes, spanning 16 edit types in 4 semantic categories, with captions at 3 verbosity levels. Each (base, var) image pair differs by exactly one localized edit while the rest of the scene is held constant, isolating compositional failures from broader scene-composition confounds.
Files
variations.parquet — 1,776 rows. Each row… See the full description on the dataset page: https://huggingface.co/datasets/CompoSET/CompoSET.heb-composed-5style-hh104-hh1-hh2-hh3-hh106-BYheb-composed-6style-hh104-hh1-hh2-hh3-hh105-hh106heb-bigram-composedheb-composed-2style-hh104-hh1heb-composed-4style-hh104-hh1-hh2-hh3-BY-450kheb-composed-2style-hh104-hh1-BYheb-composed-4style-hh104-hh1-hh2-hh3-BY-150kheb-composed-3style-hh104-hh1-hh2-BYheb-composed-4style-hh104-hh1-hh2-hh3-BYheb-composed-1style-hh104-BYeasyr1-10k-hard-qwen7b-easy-gta1-composed-20-dual-5-montage-4MP
easyr1-10k-hard-qwen7b-easy-gta1-composed-20-dual-5-montage-4MP
This dataset was generated using the enhanced EasyR1 grounding dataset pipeline with composition capabilities.
Generation Details
Generated on: 2025-08-24 11:02:10 UTC
Script: push_easyr1_composed_to_hf.py
Data directory: /lustre/fs12/portfolios/nvr/projects/nvr_lacr_llm/users/aawadalla/LLaMA-Factory/data
Parameters Used
Maximum samples: 10000
Image resize (max megapixels): 4.0 MP
Minimum native… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-10k-hard-qwen7b-easy-gta1-composed-20-dual-5-montage-4MP.easyr1-v2-pro-apps-plus-icon-data-from-yt-1k-4MP-composed
easyr1-v2-pro-apps-plus-icon-data-from-yt-1k-4MP-composed
This dataset was generated using the enhanced EasyR1 grounding dataset pipeline with composition capabilities.
Generation Details
Generated on: 2025-09-04 00:43:47 UTC
Script: push_easyr1_composed_to_hf.py
Data directory: /p/project1/synthlaion/awadalla1/datasets
Parameters Used
Maximum samples: 1000
Image resize (max megapixels): 4.0 MP
Minimum native image resolution: 0.0 MP
Prompt format: gta1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-v2-pro-apps-plus-icon-data-from-yt-1k-4MP-composed.heb-composed-4style-hh104-hh1-hh2-hh3-BY-50kTAL_OCR_Composed_37K
Dataset Information
TAL_OCR_Composed_37K is a composed set of TAL OCR datasets. The original images and labels are downloaded from TAL website (https://ai.100tal.com/dataset), includes mainly
K12 Handwritten Chinese texts, English texts, Math formulas.
Printed K12 materials.
Images are tiled randomly to have a more compact view. There are total 32K images and text pairs after processing:
TAL_OCR_CHN/composed (645 images)
TAL_OCR_ENG/composed (399 images)
TAL_OCR_MATH/composed… See the full description on the dataset page: https://huggingface.co/datasets/boydcheung/TAL_OCR_Composed_37K.heb-composed-4style-hh104-hh1-hh2-hh3-BY-350kheb-composed-1style-hh104
