8Planetterraforming/multimodal-constraint-evals
Multimodal Constraint Evaluation Dataset This dataset contains human-designed evaluation cases for multimodal image generation models. Purpose The goal of this dataset is to expose repeatable failure modes related to: object counting under strict constraints, loss of uniqueness across generated entities, layout and panel consistency, multi-step and multi-surface reasoning, planning vs rendering behavior in single-pass generation. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/8Planetterraforming/multimodal-constraint-evals.
Multimodal Constraint Evaluation Dataset
This dataset contains human-designed evaluation cases for multimodal image generation models.
Purpose
The goal of this dataset is to expose repeatable failure modes related to:
- object counting under strict constraints,
- loss of uniqueness across generated entities,
- layout and panel consistency,
- multi-step and multi-surface reasoning,
- planning vs rendering behavior in single-pass generation.
Dataset Structure
The dataset consists of prompt-based evaluation cases with:
- prompt description,
- expected behavior,
- commonly observed failure patterns.
No images are included by design. The dataset evaluates model behavior, not visual style.
Intended Use
- Human-in-the-loop model evaluation
- Multimodal reasoning benchmarking
- Analysis of structural reliability in generative models
License
This dataset is licensed under CC BY 4.0.
