figure
Datasets
All datasets matching “figure”figure2data-database-v2
figure2data databasev2 — 40K v6 合成科研图表数据集(最终交付版)
生成日期:2026-09-17 生成器:generator 1.6.0 / dataset_generation_revision v6(含两次 hotfix)
规模:40,000 样本(10 图族 / 33 亚型;area、matrix 冻结不生成,set_relation 已删除)
目录结构
databasev2/
├── figure2data.sqlite3 # 主数据库(2.0 GB:40,000 samples / 280,000 documents)
├── schema/ # sqlite schema
├── shards/ # 数据资产(shard = (样本序号-1)//1000)
│ └── shard_000 .. shard_039/
│ ├── images/ # PNG… See the full description on the dataset page: https://huggingface.co/datasets/ZZoutian/figure2data-database-v2.figureqafiguresFigureBench
FigureBench
The first large-scale benchmark for generating scientific illustrations from long-form scientific texts.
Paper | Code
Overview
FigureBench is curated to encompass a wide array of document types, including research papers, surveys, technical blogs, and textbooks, establishing a challenging and diverse testbed to spur research in automatic scientific illustration generation.
This dataset contains:
Development Set (dev): 3,000 samples with conversation-format… See the full description on the dataset page: https://huggingface.co/datasets/WestlakeNLP/FigureBench.dancing-stick-figures
Dancing Stick Figures — v0.2
A small, fully-labelled synthetic video dataset for learning (and teaching) video diffusion on one consumer GPU.
1,340 clips · 6 s @ 20 fps · 128×128 RGBA · 482,400 frames · 134 text prompts × 10 seeds × 3 cameras ·
every frame carries the 3D skeleton, camera and G-buffer (depth, normals, part segmentation) that produced it.
Think of it as an MNIST for video generation: small enough that a 64² video diffusion model trains from scratch
in a few… See the full description on the dataset page: https://huggingface.co/datasets/sprited/dancing-stick-figures.interconnects-figures
