datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GenPoster100K
Dataset Card for GenPoster100K
Dataset Summary
GenPoster-100K is a large-scale dataset for content-aware graphic layout generation introduced in the SEGA paper.
The paper describes it as a high-quality poster dataset with layer-parseable source materials and rich metadata.
This repository provides a Hugging Face datasets loader implementation that reads the source release (BruceW91/GenPoster-100K) and exposes normalized examples with:
poster background image… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/GenPoster100K.PKU-PosterLayout
Dataset Card for PKU-PosterLayout
Dataset Summary
PKU-PosterLayout is a content-aware visual-textual poster layout benchmark released with PosterLayout: A New Benchmark and Approach for Content-aware Visual-Textual Presentation Layout. The paper defines the task as arranging predefined text, logo, and underlay elements on a non-empty poster canvas while considering both inter-element and inter-layer relationships. The original benchmark contains 9,974… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PKU-PosterLayout.PubLayNet
Dataset Card for PubLayNet
Dataset Summary
PubLayNet is a large document layout analysis dataset built by automatically matching XML representations and PDF content from more than one million PubMed Central Open Access articles. It contains more than 360,000 document images with COCO-style annotations for common layout elements such as text, title, list, table, and figure regions.
Supported Tasks and Leaderboards
The dataset supports document… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PubLayNet.CreativePSD
Dataset Card for CreativePSD
Dataset Summary
CreativePSD is the PSD-derived graphic design dataset released with PSDesigner. Each example is a poster archive containing PSD tree text, structured layer metadata, tool-call trajectories, source image resources, and stepwise rendered images.
This loader keeps the contents of each poster_*.zip archive: all metadata text/JSON files, all raw_resource images, all rendering_imgs images, and a manifest of every member in… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CreativePSD.Rico
Dataset Card for Rico
Dataset Summary
Rico is a mobile app UI dataset for building data-driven design applications. The original dataset mines Android apps at runtime and exposes visual, textual, structural, and interactive design properties from more than 9.3k apps across 27 categories and more than 66k unique UI screens. This packaging provides metadata, screenshots, view hierarchies, and semantic annotations as separate configs.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/Rico.Designed-Vocalizations-Dataset
Designed Vocalizations Dataset
Paper · Demo & audio samples
The Designed Vocalizations Dataset supports voice conversion for designed vocalizations
— monster growls, robotic voices, and other sound-designed timbres — an area left
underexplored by benchmarks that focus on natural human speech. It curates diverse raw vocal
sources (speech and animal / non-linguistic sounds) and applies professional vocal-effects
processing to produce corresponding effect-modified variants. A… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/Designed-Vocalizations-Dataset.PrismLayersPro
PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models
We introduce PrismLayersPro, a 20K high-quality multi-layer transparent image dataset with rewritten style captions and human filtering.
PrismLayersPro is curated from our 200K dataset, PrismLayers, generated via MultiLayerFLUX.
Dataset Structure
📑 Dataset Splits (by Style)
The PrismLayersPro dataset is divided into 21 splits based on visual style categories.Each… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PrismLayersPro.CGL-Dataset
Dataset Card for CGL-Dataset
Dataset Summary
CGL-Dataset is a poster layout dataset released with Composition-aware Graphic Layout GAN for Visual-Textual Presentation Designs. The paper studies layout generation for a given image, emphasizing that both global semantics and spatial image composition affect where graphic elements should be placed. The original dataset contains 60,548 advertising posters with annotated layout information.
Supported… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CGL-Dataset.CGL-Dataset-v2
Dataset Card for CGL-Dataset v2
Dataset Summary
CGL-Dataset v2 is an advertising-poster layout dataset released with Relation-Aware Diffusion Model for Controllable Poster Layout Generation. The paper argues that poster layouts should account for both visual-textual relationships and geometry relationships between elements. This version extends CGL-Dataset with richer element annotations, text annotations, and text features for controllable poster layout… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CGL-Dataset-v2.Magazine
Dataset Card for Magazine
Dataset Summary
Magazine is a magazine layout dataset released with Content-aware Generative Modeling of Graphic Design Layouts. The paper studies graphic layout generation conditioned on visual and textual content and introduces a large-scale magazine layout dataset with fine-grained layout annotations and keyword labels.
Supported Tasks and Leaderboards
The dataset supports content-aware layout generation, graphic… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/Magazine.mobile-ui-design
Dataset: Mobile UI Design Detection
Introduction
This dataset is designed for object detection tasks with a focus on detecting elements in mobile UI designs. The targeted objects include text, images, and groups. The dataset contains images and object detection boxes, including class labels and location information.
Dataset Content
Load the dataset and take a look at an example:
>>> from datasets import load_dataset
>>>> ds =… See the full description on the dataset page: https://huggingface.co/datasets/mrtoy/mobile-ui-design.Design2Code-hfThis dataset consists of 484 webpages from the C4 validation set, serving the purpose of testing multimodal LLMs on converting visual designs into code implementations.
See the dataset in the raw files format here.
Note that all images in these webpages are replaced by a placeholder image (rick.jpg)
Please refer to our project page and our paper for more information.
protein_designdesign-bench
SciModelingBench Design-Bench Data
Canonical, provenance-tracked observations for scientific modeling and design Tasks.
GitHub
·
Python Package
·
Documentation
·
Organization
This repository stores the scientific observation layer used by the
SciModelingBench Design-Bench suite. The Python package supplies validators,
Agent-visible Protocols, trusted Objectives, submission contracts, and Task
metrics. Data and evaluation logic… See the full description on the dataset page: https://huggingface.co/datasets/sci-modeling-bench/design-bench.POSTAPosterArt
Dataset Card for POSTA-PosterArt
Dataset Summary
POSTA-PosterArt is the dataset introduced with POSTA, a framework for customized artistic poster generation. It contains two subsets:
PosterArt-Design: poster backgrounds with professional layout and typography annotations extracted from PSD files.
PosterArt-Text: poster title regions with artistic text captions, masks, and single-region mask images for text stylization and segmentation.
The dataset supports… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/POSTAPosterArt.PittImageVideoAdsDataset
Dataset Card for PittImageVideoAdsDataset
Dataset Summary
PittImageVideoAdsDataset is the image and video advertisement dataset released with Automatic Understanding of Image and Video Advertisements. The paper reports 64,832 image advertisements and 3,477 YouTube advertisement videos, with human annotations for topics, sentiments, slogans, persuasive strategies, symbolic references, and action/reason Q/A. This Hugging Face version exposes the public annotation… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PittImageVideoAdsDataset.protein-ligand-design
🧪 Protein-Ligand Design Gym — Team JAMMY
poolside Laguna Hackathon submission. A tool-use reinforcement-learning
environment that teaches an LLM to reason like a bench computational chemist /
protein engineer — by measuring, not guessing.
The problem
Proteins are the molecular machines inside living cells, each built from a long
string of amino-acid "letters". Ligands are the small molecules — most drugs
among them — that bind to a protein to switch it on or… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/protein-ligand-design.LICA
Dataset Card for LICA
Dataset Summary
LICA (Layered Image Composition Annotations for Graphic Design Research) is a dataset of graphic design layouts with rendered compositions, component-level layout specifications, and natural-language annotations. The public sample groups layouts by template and includes per-layout metadata, rendered PNG or MP4 files, layout JSON, per-layout annotations, and template-level annotations.
The dataset is intended for research on… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/LICA.photoshop-creative-design-trajectories
Creative-Design Computer-Use Trajectories (Preview)
A preview release of long-horizon computer-use agent trajectories from professional creative-design work in Adobe Photoshop and the browser (building advertising campaigns and fashion-editorial assets). Each step pairs a screenshot with a structured action and a first-person thought transcribed from the expert's spoken narration as they worked, so the step-level reasoning is grounded in real human intent rather than synthesized… See the full description on the dataset page: https://huggingface.co/datasets/contralabs/photoshop-creative-design-trajectories.FR5_task1_move_the_brown_colored_glass_bottle_to_the_designated_location_200This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "fairino_follower",
"total_episodes": 200,
"total_frames": 119190,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/coport-uni/FR5_task1_move_the_brown_colored_glass_bottle_to_the_designated_location_200.mobile-ui-design
Dataset: Mobile UI Design Detection
Introduction
This dataset is designed for object detection tasks with a focus on detecting elements in mobile UI designs. The targeted objects include text, images, and groups. The dataset contains images and object detection boxes, including class labels and location information.
Dataset Content
Load the dataset and take a look at an example:
>>> from datasets import load_dataset
>>>> ds =… See the full description on the dataset page: https://huggingface.co/datasets/merve/mobile-ui-design.scry-design-diff-eval
Scry Design Diff Eval
Measuring VLMs as Mobile UI Regression Reviewers
Scry Design Diff Eval is a benchmark built to evaluate vision-language models on
mobile UI diff review. Each example pairs a reference mobile screenshot
(image_a) with a generated implementation screenshot (image_b) and carries
human-drawn selection boxes plus explicit defect tags. A model must return a
structured list of UI defects — tags and normalized boxes — not a prose
description.
📄 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/Scrymore/scry-design-diff-eval.semiconductor-rtl-verilog-chip-design-2026
⚡ Complete 2026 Semiconductor & RTL/Verilog Chip Design SFT & DPO Suite
An industry-first, production-grade reasoning and alignment corpus specifically engineered for fine-tuning Large Language Models on synthesizable SystemVerilog, FPGA/ASIC hardware design, and silicon signoff verification.
This release provides 1,000 verified preview pairs (from the complete 10,000 SFT & 2,500 DPO master suite) spanning 20 mission-critical silicon IP architectures, audited against IEEE… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/semiconductor-rtl-verilog-chip-design-2026.designThe evaluation code is implemented based on MTEB framework and avaliable in https://github.com/rebeccaz4/MRMR.
creative-ad-design-dataset
Ad Creative Design Dataset (Preview)
This is a preview release of finished ad creatives produced by professional designers working from complete brand briefs. Each of the 35 rows is one fictional consumer brand: the designer received the brand's written guidelines, its logo, and a product photograph, and delivered a square social ad creative that was reviewed and approved. The row carries all three images alongside the brief broken out into structured fields: personality, color… See the full description on the dataset page: https://huggingface.co/datasets/contralabs/creative-ad-design-dataset.PosterErase
Dataset Card for PosterErase
Dataset Summary
PosterErase is a poster text-erasing dataset released with Self-supervised Text Erasing with Controllable Image Synthesis. It contains high-resolution poster images with text regions and structured annotations for text-erasing research.
This Hugging Face version exposes the original train, validation, and test splits as parquet files. The validation and test splits include ground-truth erased poster images; the… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PosterErase.engineering_design_factsDataset Copyright - L. Siddharth, Singapore University of Technology and Design, Singapore.
The dataset includes 375,084 example sentences (187200 positive, 187884 negative), each including a pair of entities and the engineering design relation between these.
The dataset was manually constructed using sentences in 4,205 patents granted by USPTO, stratified according to 130 classes.
The dataset is used to train token classification and Seq2Seq transformer models to populate explicit engineering… See the full description on the dataset page: https://huggingface.co/datasets/siddharthl1293/engineering_design_facts.Desigen
Dataset Card for Desigen
Dataset Summary
Desigen contains web advertisement design data with background images, content prompts, layout element annotations, and design canvas sizes. This loader reads the parquet shards hosted at creative-graphic-design/Desigen.
Dataset Structure
Data Fields
image: Background or rendered advertisement image.
prompt: Text prompt associated with the background image.
region: Image region boxes.… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/Desigen.GraphicDesignEvaluation
Dataset Card for GraphicDesignEvaluation
Dataset Summary
GraphicDesignEvaluation is a human-rated benchmark released with Can GPTs Evaluate Graphic Design Based on Design Principles?. The paper compares GPT-based evaluation and heuristic metrics against human ratings for three representative design principles: alignment, overlap, and white space. The dataset contains graphic banner designs curated from an online service, perturbed low-quality variants, and human… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/GraphicDesignEvaluation.DesignBench
Dataset Card for DesignBench
Dataset Summary
DesignBench is a multi-framework, multi-task benchmark for evaluating MLLM-based front-end engineering. The paper targets limitations of prior UI code generation benchmarks by covering React, Vue, Angular, and vanilla HTML/CSS, and by evaluating generation, edit, and repair workflows. The full benchmark contains 900 webpage samples spanning multiple topics, edit types, and issue categories.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/DesignBench.
