datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
claude-protein-binder-design
Claude protein binder design — data release v1.0
1,440 de novo miniprotein binders (50 to 120 residues) against 16 targets, designed by two Claude models operating as autonomous protein-design agents (Mythos Preview, 900 designs; Opus 4.8, 540 designs) and characterized at two contract research organizations, Adaptyv Bio (cell-free expression; SPR/BLI kinetics with the design immobilized) and Twist Bioscience (Fc-fusion expression; capture SPR with a six-point antigen… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/claude-protein-binder-design.GenPoster100K
Dataset Card for GenPoster100K
Dataset Summary
GenPoster-100K is a large-scale dataset for content-aware graphic layout generation introduced in the SEGA paper.
The paper describes it as a high-quality poster dataset with layer-parseable source materials and rich metadata.
This repository provides a Hugging Face datasets loader implementation that reads the source release (BruceW91/GenPoster-100K) and exposes normalized examples with:
poster background image… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/GenPoster100K.PKU-PosterLayout
Dataset Card for PKU-PosterLayout
Dataset Summary
PKU-PosterLayout is a content-aware visual-textual poster layout benchmark released with PosterLayout: A New Benchmark and Approach for Content-aware Visual-Textual Presentation Layout. The paper defines the task as arranging predefined text, logo, and underlay elements on a non-empty poster canvas while considering both inter-element and inter-layer relationships. The original benchmark contains 9,974… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PKU-PosterLayout.design_bench_dataPubLayNet
Dataset Card for PubLayNet
Dataset Summary
PubLayNet is a large document layout analysis dataset built by automatically matching XML representations and PDF content from more than one million PubMed Central Open Access articles. It contains more than 360,000 document images with COCO-style annotations for common layout elements such as text, title, list, table, and figure regions.
Supported Tasks and Leaderboards
The dataset supports document… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PubLayNet.Design2CodeThis dataset consists of 484 webpages from the C4 validation set, serving the purpose of testing multimodal LLMs on converting visual designs into code implementations.
Each example is a pair of source HTML and screenshot ({id}.html and {id}.png).
See the dataset in the huggingface format here.
Note that all images in these webpages are replaced by a placeholder image (rick.jpg)
Please refer to our project page and our paper for more information.
Example Usage
For example, you… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/Design2Code.DesignBenchResultsCreativePSD
Dataset Card for CreativePSD
Dataset Summary
CreativePSD is the PSD-derived graphic design dataset released with PSDesigner. Each example is a poster archive containing PSD tree text, structured layer metadata, tool-call trajectories, source image resources, and stepwise rendered images.
This loader keeps the contents of each poster_*.zip archive: all metadata text/JSON files, all raw_resource images, all rendering_imgs images, and a manifest of every member in… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CreativePSD.DesignBenchRico
Dataset Card for Rico
Dataset Summary
Rico is a mobile app UI dataset for building data-driven design applications. The original dataset mines Android apps at runtime and exposes visual, textual, structural, and interactive design properties from more than 9.3k apps across 27 categories and more than 66k unique UI screens. This packaging provides metadata, screenshots, view hierarchies, and semantic annotations as separate configs.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/Rico.Chip_Design
AURA-1: designing an edge-AI chip for a neckband headphone
Working files from a solo, from-scratch attempt to specify and prototype an
edge-AI inference SoC for wireless neckband headphones: a resident,
speech-native conversational model on the device, the cloud called as a tool for
facts. The novel block is the NPU; Wi-Fi, Bluetooth, codec, ANC and PMU are
sourced, not designed.
The author is a process engineer learning chip design. The archive is
deliberately complete: the… See the full description on the dataset page: https://huggingface.co/datasets/sangramrout/Chip_Design.fashion_designdesign-patents-not-in-impact
US Design Patents Not Included in IMPACT (2008-2026)
Original drawing images (TIFF) and grant full-text XML for 165,917 US design patents that are
absent from the AI4Patents/IMPACT dataset.
IMPACT covers 2007-2022 and contains 434,498 rows. This dataset supplies the design patents that
IMPACT does not have: 161,093 patents granted in 2023-2026, which are outside IMPACT's period,
plus 4,824 patents from years IMPACT does cover but did not include. There is no patent
overlap with… See the full description on the dataset page: https://huggingface.co/datasets/SoichiOnozuka/design-patents-not-in-impact.fashion_design_qaDesigned-Vocalizations-Dataset
Designed Vocalizations Dataset
Paper · Demo & audio samples
The Designed Vocalizations Dataset supports voice conversion for designed vocalizations
— monster growls, robotic voices, and other sound-designed timbres — an area left
underexplored by benchmarks that focus on natural human speech. It curates diverse raw vocal
sources (speech and animal / non-linguistic sounds) and applies professional vocal-effects
processing to produce corresponding effect-modified variants. A… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/Designed-Vocalizations-Dataset.design_bench_dataPrismLayersPro
PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models
We introduce PrismLayersPro, a 20K high-quality multi-layer transparent image dataset with rewritten style captions and human filtering.
PrismLayersPro is curated from our 200K dataset, PrismLayers, generated via MultiLayerFLUX.
Dataset Structure
📑 Dataset Splits (by Style)
The PrismLayersPro dataset is divided into 21 splits based on visual style categories.Each… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PrismLayersPro.CGL-Dataset
Dataset Card for CGL-Dataset
Dataset Summary
CGL-Dataset is a poster layout dataset released with Composition-aware Graphic Layout GAN for Visual-Textual Presentation Designs. The paper studies layout generation for a given image, emphasizing that both global semantics and spatial image composition affect where graphic elements should be placed. The original dataset contains 60,548 advertising posters with annotated layout information.
Supported… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CGL-Dataset.CGL-Dataset-v2
Dataset Card for CGL-Dataset v2
Dataset Summary
CGL-Dataset v2 is an advertising-poster layout dataset released with Relation-Aware Diffusion Model for Controllable Poster Layout Generation. The paper argues that poster layouts should account for both visual-textual relationships and geometry relationships between elements. This version extends CGL-Dataset with richer element annotations, text annotations, and text features for controllable poster layout… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CGL-Dataset-v2.bricklink_lego_designNovel_NLRP3_Inhibitor_Designs
Novel NLRP3 Inhibitors — GA-II Designed Ligand–Receptor Complexes
Why this target matters. NLRP3 sits upstream of IL-1β in gout, cardiovascular, metabolic and neurodegenerative disease, yet after two decades of effort no small-molecule NLRP3 inhibitor has reached approval.
176 small molecules generated de novo by the Technetium TC-43.ai engine (GA-II), conditioned on the allosteric site of the NLRP3 inflammasome NACHT module.
Each molecule was constructed against this pocket… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/Novel_NLRP3_Inhibitor_Designs.Magazine
Dataset Card for Magazine
Dataset Summary
Magazine is a magazine layout dataset released with Content-aware Generative Modeling of Graphic Design Layouts. The paper studies graphic layout generation conditioned on visual and textual content and introduces a large-scale magazine layout dataset with fine-grained layout annotations and keyword labels.
Supported Tasks and Leaderboards
The dataset supports content-aware layout generation, graphic… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/Magazine.game-design-pattern-core-collection
Game Design Patterns Dataset
Original Source Attribution
This dataset is derived from the work of Staffan Björk and Jussi Holopainen. The original content comes from "Patterns in Game Design," published by Charles River Media in 2005.
Original authors: Jussi Kuittinen, Staffan Björk and Jussi Holopainen
Original format: HTML documents publicly available at: https://www.researchgate.net/publication/379683418_collection.zip
Citation: Bjork, S., & Holopainen, J. (2005).… See the full description on the dataset page: https://huggingface.co/datasets/HughXuechen/game-design-pattern-core-collection.Design2Code-HARDThis dataset consists of 80 extra difficult webpages from Github Pages, which challenges SoTA multimodal LLMs on converting visual designs into code implementations.
Each example is a pair of source HTML and screenshot ({id}.html and {id}.png).
See the "easy" version of the Design2Code testset here
Note that all images in these webpages are replaced by a placeholder image (rick.jpg)
Please refer to our project page and our paper for more information.
mobile-ui-design
Dataset: Mobile UI Design Detection
Introduction
This dataset is designed for object detection tasks with a focus on detecting elements in mobile UI designs. The targeted objects include text, images, and groups. The dataset contains images and object detection boxes, including class labels and location information.
Dataset Content
Load the dataset and take a look at an example:
>>> from datasets import load_dataset
>>>> ds =… See the full description on the dataset page: https://huggingface.co/datasets/mrtoy/mobile-ui-design.Design2Code-hfThis dataset consists of 484 webpages from the C4 validation set, serving the purpose of testing multimodal LLMs on converting visual designs into code implementations.
See the dataset in the raw files format here.
Note that all images in these webpages are replaced by a placeholder image (rick.jpg)
Please refer to our project page and our paper for more information.
PittImageVideoAdsDataset
Dataset Card for PittImageVideoAdsDataset
Dataset Summary
PittImageVideoAdsDataset is the image and video advertisement dataset released with Automatic Understanding of Image and Video Advertisements. The paper reports 64,832 image advertisements and 3,477 YouTube advertisement videos, with human annotations for topics, sentiments, slogans, persuasive strategies, symbolic references, and action/reason Q/A. This Hugging Face version exposes the public annotation… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PittImageVideoAdsDataset.claude-protein-binder-design-dgui-corpus
Claude protein binder design — data release v1.0
1,440 de novo miniprotein binders (50 to 120 residues) against 16 targets, designed by two Claude models operating as autonomous protein-design agents (Mythos Preview, 900 designs; Opus 4.8, 540 designs) and characterized at two contract research organizations, Adaptyv Bio (cell-free expression; SPR/BLI kinetics with the design immobilized) and Twist Bioscience (Fc-fusion expression; capture SPR with a six-point antigen… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/claude-protein-binder-design-dgui-corpus.DesignVFR
DesignVFR
Towards Universal Open-Set Visual Font Recognition via Augmented Synthetic Similarity · CVPR 2026 Findings
DesignVFR is the first large-scale dataset for universal open-set Visual Font Recognition (VFR). While prior VFR work is limited to closed-set classification on isolated character-level grayscale images, DesignVFR covers font recognition in real-world universal scenarios — sentences, complex backgrounds, and artistic effects across posters, films… See the full description on the dataset page: https://huggingface.co/datasets/Tunanzzz/DesignVFR.PosterIQ
Dataset Card for PosterIQ
Dataset Summary
PosterIQ is the poster design benchmark released with PosterIQ: A Design Perspective Benchmark for Poster Understanding and Generation. It contains task-level evaluation data for poster understanding and poster generation from a design perspective, including typography, layout, OCR, composition, style, empty-space use, and design intention.
This Hugging Face loader exposes the upstream release as 24 task-level… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PosterIQ.
