datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wds_imagenet_sketchimagenet_sketchImageNet-Sketch data set consists of 50000 images, 50 images for each of the 1000 ImageNet classes.
We construct the data set with Google Image queries "sketch of __", where __ is the standard class name.
We only search within the "black and white" color scheme. We initially query 100 images for every class,
and then manually clean the pulled images by deleting the irrelevant images and images that are for similar
but different classes. For some classes, there are less than 50 images after manually cleaning, and then we
augment the data set by flipping and rotating the images.imagenet-sketch-datawds_imagenet_sketch_test
ImageNet-Sketch (Test set only)
Original paper: Learning Robust Global Representations by Penalizing Local Predictive Power
Homepage: https://github.com/HaohanWang/ImageNet-Sketch
Bibtex:
@inproceedings{wang2019learning,
title={Learning Robust Global Representations by Penalizing Local Predictive Power},
author={Wang, Haohan and Ge, Songwei and Lipton, Zachary and Xing, Eric P},
booktitle={Advances in Neural Information Processing Systems}… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_imagenet_sketch_test.trellis500k-sketchfab-archiveswhest-p2-bakev2-d8b-sketch-g00
whest-p2-bakev2 — cumulant sketches of 16×1024 ReLU MLPs (round 1, 2026-09-13)
Monte-Carlo cumulants of the pre-/post-activations of 1024-wide, 16-layer ReLU MLPs under
standard-normal inputs, stored as Ω-sketches (every net) plus a few dense n×n blocks
(validation nets). Companion code: ap_p2_bakev2_schema.py (seeds, Ω generator,
conversions, loader), ap_p2_bakev2.py (bake), ap_p2_bakev2_check.py (validation).
Schema version bakev2-r1-2026-09-13.
bake
family / nets
tier… See the full description on the dataset page: https://huggingface.co/datasets/keenanpepper/whest-p2-bakev2-d8b-sketch-g00.sketchvlm-physics-ball-drop
SketchVLM: Physics Ball Drop Dataset
This dataset is part of the SketchVLM project, introduced in the paper: SketchVLM: Vision language models can annotate images to explain thoughts and guide users.
Project Page | GitHub | Interactive Demo
Description
SketchVLM is a training-free, model-agnostic framework that enables vision-language models (VLMs) to produce non-destructive, editable SVG overlays on input images to visually explain their reasoning.
The Physics Ball… See the full description on the dataset page: https://huggingface.co/datasets/loganbolton/sketchvlm-physics-ball-drop.splataverse-sketchfab
Splataverse: Sketchfab
imagenet_sketch
Dataset Card for ImageNet-Sketch
This dataset was duplicated from songweig/imagenet_sketch so I could create validation splits and use parquets for caching the dataset instead of requiring python code execution.
The dataset was replicated using the following script:
from datasets import load_dataset, DatasetDict
import os
def main():
ds = load_dataset("songweig/imagenet_sketch", split="train").shuffle(seed=42)
ds_train = ds.select(range(40_000))
ds_validate =… See the full description on the dataset page: https://huggingface.co/datasets/vaughankraska/imagenet_sketch.QuickdrawHDSketch2CodeThe Sketch2Code dataset consists of 731 human-drawn sketches paired with 484 real-world webpages from the Design2Code dataset, serving to benchmark Vision-Language Models (VLMs) on converting rudimentary sketches into web design prototypes.
Each example consists of a pair of source HTML and rendered webpage screenshot (stored in webpages/ directory under name {webpage_id}.html and {webpage_id}.png), as well as 1 to 3 sketches drawn by human annotators (stored in sketches/ directory under name… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/Sketch2Code.ckasketch-sketches
ckasketch sketches — multi-method static + activation
145 sketches (48.2 GB total) for HuggingFace models, generated
by ckasketch. 108 of 145 carry
all five static-mode methods (CKA, SVD, SVD-MP, Eigen, SRHT) plus
activation arrays captured against the
ckasketch v1 text calibration corpus
(frozen 2026-05-17, hash cbd6a314d904842e..., 1053 items); the remaining
37 carry a partial subset (commonly just cka+svd-mp,
ckasketch's current generate default) and/or lack activation. Check… See the full description on the dataset page: https://huggingface.co/datasets/marcjon/ckasketch-sketches.ImageNet-Sketchsketch-scene
Dataset Card for Sketch Scene Descriptions
Dataset used to train Sketch Scene text to image model
We advance sketch research to scenes with the first dataset of freehand scene sketches, FS-COCO. With practical applications in mind, we collect sketches that convey well scene content but can be sketched within a few minutes by a person with any sketching skills. Our dataset comprises around 10,000 freehand scene vector sketches with per-point space-time information by 100 non-expert… See the full description on the dataset page: https://huggingface.co/datasets/zoheb/sketch-scene.wds_imagenet_sketchobjaversexl_sketchfabsketchybusinesssketch-generation-dissertation
High-Fidelity Image Generation from Arbitrary Abstract Sketches
MSc Dissertation Research ProjectAuthor: Sumukh MaratheSupervisor: Prof. Yi-Zhe Song (SketchX Group, University of Surrey)Date: 2025
Project Overview
This repository contains the complete implementation, documentation, and results for the dissertation project on sketch-conditioned image generation with abstraction-aware adaptive conditioning.
Core Question: How can sketch-conditioned image generation be… See the full description on the dataset page: https://huggingface.co/datasets/sumukhmarathe/sketch-generation-dissertation.FreeCAD_Sketches
🧩 FreeCAD Sketch Python Dataset
This dataset contains 3,000 Python files that define parametric FreeCAD sketches. Each file represents a geometric sketch scripted in Python and is ideal for training large language models (LLMs) or fine-tuning code models to understand and generate CAD-based geometric design code.
📁 Dataset Overview
Number of files: ~3,000
Format: Python (.py)
Language: English (en)
License: Apache 2.0
Purpose: Training LLMs for text-to-CAD and… See the full description on the dataset page: https://huggingface.co/datasets/Yas1n/FreeCAD_Sketches.Imagenet_Sketchimagenet-sketchImageNet-Sketch data set consists of 50000 images, 50 images for each of the 1000 ImageNet classes.
We construct the data set with Google Image queries "sketch of __", where __ is the standard class name.
We only search within the "black and white" color scheme. We initially query 100 images for every class,
and then manually clean the pulled images by deleting the irrelevant images and images that are for similar
but different classes. For some classes, there are less than 50 images after manually cleaning, and then we
augment the data set by flipping and rotating the images.objaversexl_sketchfab_pmap1ObjaverseXL_sketchfab_v2objaversexl_sketchfab_pmapsketch_libero_distillThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 182,
"total_frames": 22629,
"total_tasks": 10,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:182"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ductaingn/sketch_libero_distill.FreeCAD_Sketches_Pics
🧩 FreeCAD Sketch Python Dataset
This dataset contains approximately 1,000 Python files and their corresponding images defining parametric FreeCAD sketches. Each python file represents a geometric sketch scripted using the FreeCAD Python API.
The dataset is intended for training and fine-tuning large language models (LLMs) and vision large language models (vLLMs) and other AI systems to understand and generate CAD-based parametric geometry code, particularly for text-to-CAD and… See the full description on the dataset page: https://huggingface.co/datasets/Yas1n/FreeCAD_Sketches_Pics.genshin_impact_CHONGYUN_Paints_UNDO_Sketch_UnChunked
SketchGraphsspatial_sketchtikz-sketch-splits
tikz-sketch-splits
Synthetic-sketch data augmentation splits generated from
loss-boss/tikz-train, for
sketch-to-TikZ model training/evaluation.
Each row in loss-boss/tikz-train has two source variants:
with_text (image_with_text/code_with_text/llm_description_with_text)
without_text (image_without_text_full/code_without_text_full/llm_description_without_text_full only populated for a minority of rows)
For every row processed here, one variant was picked at random (falling back… See the full description on the dataset page: https://huggingface.co/datasets/03kiko/tikz-sketch-splits.
