datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
game-design-pattern-core-collection
Game Design Patterns Dataset
Original Source Attribution
This dataset is derived from the work of Staffan Björk and Jussi Holopainen. The original content comes from "Patterns in Game Design," published by Charles River Media in 2005.
Original authors: Jussi Kuittinen, Staffan Björk and Jussi Holopainen
Original format: HTML documents publicly available at: https://www.researchgate.net/publication/379683418_collection.zip
Citation: Bjork, S., & Holopainen, J. (2005).… See the full description on the dataset page: https://huggingface.co/datasets/HughXuechen/game-design-pattern-core-collection.TTS-Voice-Design-Benchmark
TTS Voice Design Benchmark
🏆 Leaderboard | 🛠️ Evaluation Suite
TTS Voice Design is a high-quality benchmark of 1,000 character voice-design
tasks spanning a broad range of media genres and real-world creative use
cases. It evaluates whether a text-to-speech model can turn an open-ended
character profile into a distinctive, appropriate, and usable voice.
Unlike benchmarks built around a fixed set of speakers or isolated acoustic
attributes, this dataset covers complete… See the full description on the dataset page: https://huggingface.co/datasets/BreezeBlue/TTS-Voice-Design-Benchmark.Bio-Design-ProcessThis dataset works even though it may not be the cleanest in regards to organization. I'm working on cleaning it up for better performance, but it should still work as long as you don't overtrain on it.
2026-08-26-sonnet45-post-action-retrospection-natural-turn-design
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
date_generated
20260826_152715
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ c2fdee460e71fa28e9902edf1cc662db0d19cad8
models
per-stage models — see… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-sonnet45-post-action-retrospection-natural-turn-design.arch-design-sft
arch-design-sft: verified architecture-design SFT data
Supervised fine-tuning data for neural architecture design treated as structured graph editing. Each row pairs a natural-language design spec and a serialized starting graph with a reference action plan, and every row is re-graded by a deterministic verifier before it is written: structural blockers, parameter budgets and bands, required layer families. Ten task families across six design-from-spec and four edit-in-place… See the full description on the dataset page: https://huggingface.co/datasets/neurarch-ai/arch-design-sft.Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam
Unity Code and GPT-Generated GDD Pairs Dataset
This dataset contains paired samples of Unity game mechanic scripts and their corresponding GPT-4 generated Game Design Documents (GDDs). It is intended for training and benchmarking LLMs in game code generation from design specifications.
Format
Each entry is stored as a .jsonl file with:
"input": GPT-4 generated GDD describing a specific game and its mechanics
"output": Unity C# scripts implementing the described mechanic… See the full description on the dataset page: https://huggingface.co/datasets/AmnaHassan/Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam.am-session-sharing-design
Agent Manager session — designing session sharing
Access: public. Anyone with the link can read this trace. Public is the only mode
where the Hub's trace viewer works for every visitor regardless of account tier.
A single real Agent Manager
session, exported as a shareable trace. This is the working session in which the
session-sharing design itself was researched and written — surveying how five coding-agent
harnesses persist sessions, testing what the Hub's trace viewer can… See the full description on the dataset page: https://huggingface.co/datasets/thomwolf/am-session-sharing-design.M-DESIGN-Knowledge-Base
M-DESIGN Knowledge Base and Model Artifacts
This dataset contains the released SQLite model-performance databases and model
artifacts used by M-DESIGN, the method from "Beyond Model Base Retrieval:
Weaving Knowledge to Master Fine-grained Neural Network Design".
Contents
Each .db file has a model_records table. The first six columns encode
fine-grained neural design choices and the final two columns store the measured
score and standard deviation. Each task/dataset… See the full description on the dataset page: https://huggingface.co/datasets/jilwang804/M-DESIGN-Knowledge-Base.am-session-sharing-design-noimg
Agent Manager session — designing session sharing
Access: public. Anyone with the link can read this trace. Public is the only mode
where the Hub's trace viewer works for every visitor regardless of account tier.
A single real Agent Manager
session, exported as a shareable trace. This is the working session in which the
session-sharing design itself was researched and written — surveying how five coding-agent
harnesses persist sessions, testing what the Hub's trace viewer can… See the full description on the dataset page: https://huggingface.co/datasets/thomwolf/am-session-sharing-design-noimg.DesignAsCode-training-data
DesignAsCode Training Data
Training data for the DesignAsCode Semantic Planner.
Overview
Samples
19,479
Format
JSONL (one JSON object per line)
Size
~145 MB
Data Source
Each sample corresponds to a real graphic design from the Crello dataset. We distilled structured design semantics from each original design using GPT-4o and GPT-o3, taking the original design, its individual layers, and design metadata as input.
The distillation produces:… See the full description on the dataset page: https://huggingface.co/datasets/Tony1109/DesignAsCode-training-data.ParaSFT-designer
ParaSFT Designer
English | 中文
Overview
ParaSFT Designer is a private supervised fine-tuning dataset for ParadoxGPT-Designer-4B, the ParadoxGPT specialist model for experiment design, evidence planning, and claim-to-experiment mapping.
Designer annotation pipeline over ParaPaper context packs, covering claim-to-evidence mapping, experiment argument planning, ablation design, sufficiency critique, and interpretation boundaries.
Each example is an instruction-tuning… See the full description on the dataset page: https://huggingface.co/datasets/bhxdianzhang/ParaSFT-designer.designer-design-logics
DESIGNER: Design Logic Library [Project Page]
This repository contains a library of Mermaid-format Design Logics used in the paper DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning (ICLR 2026).
Field definitions
mermaid: Design Logic in Mermaid format, abstracted from the source question, which is a human-authored high-difficulty question.
difficulty: difficulty label of the source question
type: type label of the source question… See the full description on the dataset page: https://huggingface.co/datasets/Attention1115/designer-design-logics.nonattainment-county-designations-by-pollutant
Current EPA nonattainment county designations by pollutant
Canonical, always-current version: https://referencesource.org/nonattainment-county-designations-by-pollutant/
Machine-readable: https://referencesource.org/nonattainment-county-designations-by-pollutant/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-24
Stale after: 2027-02-20 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 463
Whether a… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nonattainment-county-designations-by-pollutant.web-design-diamond
Web Design Diamond — dataset SFT
Dataset para entrenar LLMs (chicos: 2B-8B) que, dado un pedido en lenguaje natural, generan
UNA pagina index.html autocontenida (Tailwind CSS por CDN + JS vanilla embebido) que funciona
y se ve premium, rapido. Texto -> codigo (no imagen -> codigo).
Que tiene de distinto
Razonamiento (thinking) secuencial y sin leakage: el modelo razona que va a hacer ANTES
de implementar; el thinking previo a una tool no menciona tokens… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/web-design-diamond.ptdbench-reward-design-reward-polynomial-factorization-035-dataset
PTDBench dataset snapshot: reward_polynomial_factorization_035
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-polynomial-factorization-035-dataset.design_compiler_pdf_datasetllm-system-design-guide-basicbit_new_design_1fresh-by-design-cosmetic-lexicon
Fresh by Design Cosmetic Lexicon v1.0
This dataset is a small bilingual lexicon of cosmetic freshness terms used to describe Sostenica's "Fresh by Design" / "Fresco Por Diseño" formulation philosophy. Version 1.0 is aligned with Sostenica's public Glosario de Formulación.
The dataset is intentionally small. Its purpose is not model training at scale. Its purpose is structured terminology: a machine-readable reference for a freshness-oriented cosmetic vocabulary around formulation… See the full description on the dataset page: https://huggingface.co/datasets/sostenica/fresh-by-design-cosmetic-lexicon.system_design
data ingestion product Nexus
Description
It's about the project Nexus whose task is to gather the data from all the media, video, digital channels of a company and keep ingesting it thus keeping the knowledge base embeddings up to date.
Format
This dataset is in alpaca format.
Creation Method
This dataset was created using the Easy Dataset tool.
Easy Dataset is a specialized application designed to streamline the creation of fine-tuning datasets for… See the full description on the dataset page: https://huggingface.co/datasets/SaffalPoosh/system_design.MF-Design
Data for MF-Design
This repository hosts the datasets for the paper "Repurposing AlphaFold3-like Protein Folding Models for Antibody Sequence and Structure Co-design" (MF-Design).
The main code repository can be found at MF-Design.
Directory Structure
The ./data/ directory contains all the necessary data for running and evaluating the models.
Raw and Processed Data: Includes the original data used for training in ./data/raw_data.tar.zst and its processed versions in… See the full description on the dataset page: https://huggingface.co/datasets/clorf6/MF-Design.ptdbench-reward-design-reward-difference-constraint-system-034-dataset
PTDBench dataset snapshot: reward_difference_constraint_system_034
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-difference-constraint-system-034-dataset.Experiment-Design-Sketch-Image-Classification-Dataset
Experiment Design Sketch Image Classification Dataset
In the field of industrial manufacturing, the design process often relies on a large number of design sketches that need to be quickly converted into actual engineering designs during the subsequent manufacturing stages. However, manually processing these sketches is often time-consuming and prone to errors, currently relying mainly on manual labeling and conversion by designers, which is inefficient and unstable. Existing… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Experiment-Design-Sketch-Image-Classification-Dataset.ptdbench-reward-design-reward-visible-line-038-dataset
PTDBench dataset snapshot: reward_visible_line_038
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every hydrated… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-visible-line-038-dataset.english-ui-ux-design-basics-30design_compiler_md_datasetCreative-Design-Studio-Interaction-Body-Language-Recognition-Video-Dataset
Creative Design Studio Interaction Body Language Recognition Video Dataset
In today's creative design industry, understanding the complex non-verbal communication among team members is crucial. However, existing body language recognition methods perform limitedly in dense interactive environments, struggling to accurately interpret subtle body movements and postures. This video dataset aims to tackle the technical challenges of body language analysis in creative discussions… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Creative-Design-Studio-Interaction-Body-Language-Recognition-Video-Dataset.website-hero-designptdbench-reward-design-reward-integer-programming-029-dataset
PTDBench dataset snapshot: reward_integer_programming_029
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every hydrated… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-integer-programming-029-dataset.ptdbench-reward-design-reward-prefix-product-mod-distinct-permutation-011-dataset
PTDBench dataset snapshot: reward_prefix_product_mod_distinct_permutation_011
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-prefix-product-mod-distinct-permutation-011-dataset.
