datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vargov-design-catalog
Vargov® Design Catalog — 605 lighting and decorative compositions in 8 languages
A machine-readable catalog of the full body of work of Vargov® Design, an author-driven
studio of lighting and decorative compositions founded by designer Anton Vargov (Moscow).
Every record is one composition: its identifier, category, canonical URLs, image links,
awards, links to its 3D model, and editorial copy written by the studio in eight
languages — Russian, English, German, Italian, French… See the full description on the dataset page: https://huggingface.co/datasets/vargov-design/vargov-design-catalog.semiconductor-rtl-verilog-chip-design-2026
⚡ Complete 2026 Semiconductor & RTL/Verilog Chip Design SFT & DPO Suite
An industry-first, production-grade reasoning and alignment corpus specifically engineered for fine-tuning Large Language Models on synthesizable SystemVerilog, FPGA/ASIC hardware design, and silicon signoff verification.
This release provides 1,000 verified preview pairs (from the complete 10,000 SFT & 2,500 DPO master suite) spanning 20 mission-critical silicon IP architectures, audited against IEEE… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/semiconductor-rtl-verilog-chip-design-2026.DesignBench
Dataset Card for DesignBench
Dataset Summary
DesignBench is a multi-framework, multi-task benchmark for evaluating MLLM-based front-end engineering. The paper targets limitations of prior UI code generation benchmarks by covering React, Vue, Angular, and vanilla HTML/CSS, and by evaluating generation, edit, and repair workflows. The full benchmark contains 900 webpage samples spanning multiple topics, edit types, and issue categories.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/DesignBench.arch-design-sft
arch-design-sft: verified architecture-design SFT data
Supervised fine-tuning data for neural architecture design treated as structured graph editing. Each row pairs a natural-language design spec and a serialized starting graph with a reference action plan, and every row is re-graded by a deterministic verifier before it is written: structural blockers, parameter budgets and bands, required layer families. Ten task families across six design-from-spec and four edit-in-place… See the full description on the dataset page: https://huggingface.co/datasets/neurarch-ai/arch-design-sft.ParaSFT-designer
ParaSFT Designer
English | 中文
Overview
ParaSFT Designer is a private supervised fine-tuning dataset for ParadoxGPT-Designer-4B, the ParadoxGPT specialist model for experiment design, evidence planning, and claim-to-experiment mapping.
Designer annotation pipeline over ParaPaper context packs, covering claim-to-evidence mapping, experiment argument planning, ablation design, sufficiency critique, and interpretation boundaries.
Each example is an instruction-tuning… See the full description on the dataset page: https://huggingface.co/datasets/bhxdianzhang/ParaSFT-designer.dolma3_mix-common_crawl-art_and_design-160kThe 160K subset of AllenAI's common_crawl-art_and_design Pretraining dataset split into train a valid saamples.
Train set size: 159436
Valid set size: 160
Direct usage in MLX-LM-LoRA:
python -m mlx_lm.lora \
--train \
--model Qwen/Qwen3-0.6B-Base \
--data mlx-community/dolma3_mix-common_crawl-art_and_design-160k \
--num-layers 4 \
--iters 1000 \
--batch-size 1 \
--steps-per-report 50 \
--max-seq-length 1028 \
--adapter-path path/to/adapter
Direct usage in MLX-LM:
python -m mlx_lm.lora \… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/dolma3_mix-common_crawl-art_and_design-160k.DOD-Enterprise-DevSecOps-Reference-Design-AWS-Managed-Services
DoD Enterprise DevSecOps AWS Managed Services Reference Design Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on DoD Enterprise DevSecOps Reference Design: AWS Managed Services (DoD IaC Baseline), Version 0.2, September 2021.
The source presents a draft Department of Defense reference design for implementing a DevSecOps software factory… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Enterprise-DevSecOps-Reference-Design-AWS-Managed-Services.DesignAsCode-training-data
DesignAsCode Training Data
Training data for the DesignAsCode Semantic Planner.
Overview
Samples
19,479
Format
JSONL (one JSON object per line)
Size
~145 MB
Data Source
Each sample corresponds to a real graphic design from the Crello dataset. We distilled structured design semantics from each original design using GPT-4o and GPT-o3, taking the original design, its individual layers, and design metadata as input.
The distillation produces:… See the full description on the dataset page: https://huggingface.co/datasets/Tony1109/DesignAsCode-training-data.designer-design-logics
DESIGNER: Design Logic Library [Project Page]
This repository contains a library of Mermaid-format Design Logics used in the paper DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning (ICLR 2026).
Field definitions
mermaid: Design Logic in Mermaid format, abstracted from the source question, which is a human-authored high-difficulty question.
difficulty: difficulty label of the source question
type: type label of the source question… See the full description on the dataset page: https://huggingface.co/datasets/Attention1115/designer-design-logics.web-design-diamond
Web Design Diamond — dataset SFT
Dataset para entrenar LLMs (chicos: 2B-8B) que, dado un pedido en lenguaje natural, generan
UNA pagina index.html autocontenida (Tailwind CSS por CDN + JS vanilla embebido) que funciona
y se ve premium, rapido. Texto -> codigo (no imagen -> codigo).
Que tiene de distinto
Razonamiento (thinking) secuencial y sin leakage: el modelo razona que va a hacer ANTES
de implementar; el thinking previo a una tool no menciona tokens… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/web-design-diamond.ptdbench-reward-design-reward-polynomial-factorization-035-dataset
PTDBench dataset snapshot: reward_polynomial_factorization_035
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-polynomial-factorization-035-dataset.ptdbench-reward-design-reward-difference-constraint-system-034-dataset
PTDBench dataset snapshot: reward_difference_constraint_system_034
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-difference-constraint-system-034-dataset.ptdbench-reward-design-reward-visible-line-038-dataset
PTDBench dataset snapshot: reward_visible_line_038
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every hydrated… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-visible-line-038-dataset.ptdbench-reward-design-reward-integer-programming-029-dataset
PTDBench dataset snapshot: reward_integer_programming_029
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every hydrated… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-integer-programming-029-dataset.ptdbench-reward-design-reward-min-cost-reducing-lnds-020-dataset
PTDBench dataset snapshot: reward_min_cost_reducing_lnds_020
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-min-cost-reducing-lnds-020-dataset.ptdbench-reward-design-reward-prefix-product-mod-distinct-permutation-011-dataset
PTDBench dataset snapshot: reward_prefix_product_mod_distinct_permutation_011
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-prefix-product-mod-distinct-permutation-011-dataset.ptdbench-reward-design-reward-prefix-sum-mod-distinct-permutation-010-dataset
PTDBench dataset snapshot: reward_prefix_sum_mod_distinct_permutation_010
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-prefix-sum-mod-distinct-permutation-010-dataset.ptdbench-reward-design-reward-quadratic-function-segmentation-019-dataset
PTDBench dataset snapshot: reward_quadratic_function_segmentation_019
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-quadratic-function-segmentation-019-dataset.ptdbench-reward-design-reward-max-different-group-pair-division-023-dataset
PTDBench dataset snapshot: reward_max_different_group_pair_division_023
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-max-different-group-pair-division-023-dataset.ptdbench-reward-design-reward-sat-031-dataset
PTDBench dataset snapshot: reward_sat_031
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every hydrated runtime path… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-sat-031-dataset.ptdbench-reward-design-reward-set-splitting-012-dataset
PTDBench dataset snapshot: reward_set_splitting_012
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every hydrated… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-set-splitting-012-dataset.ptdbench-reward-design-reward-topological-sort-minimal-lexicographical-order-001-dataset
PTDBench dataset snapshot: reward_topological_sort_minimal_lexicographical_order_001
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-topological-sort-minimal-lexicographical-order-001-dataset.ptdbench-reward-design-reward-land-acquisition-008-dataset
PTDBench dataset snapshot: reward_land_acquisition_008
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every hydrated… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-land-acquisition-008-dataset.ptdbench-reward-design-reward-min-pair-sum-multiplication-permutation-007-dataset
PTDBench dataset snapshot: reward_min_pair_sum_multiplication_permutation_007
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-min-pair-sum-multiplication-permutation-007-dataset.ptdbench-reward-design-reward-new-nim-game-021-dataset
PTDBench dataset snapshot: reward_new_nim_game_021
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every hydrated… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-new-nim-game-021-dataset.smolified-roast-your-design
🤏 smolified-roast-your-design
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model AitijhyaR/smolified-roast-your-design.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 17a09ac9)
Records: 93
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by AitijhyaR.
Generated via Smolify.ai.
ml-design-doc-reviewer-data
ml-system-design/ml-design-doc-reviewer-data (v1.0.0)
Evaluation artifacts for the ML Design Doc Reviewer project.
Layout
Path
Description
manifest/sample_manifest.csv
Stratified 100-case sample manifest
manifest/error_topology.csv
Controlled error taxonomy for flawed docs
raw/
Raw markdown exports, metadata sidecars, OCR image blocks
raw/images/
Downloaded article images
normalized/
Canonical 14-section ML design documents
flawed/
Normalized… See the full description on the dataset page: https://huggingface.co/datasets/ml-system-design/ml-design-doc-reviewer-data.
