datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cgaxis-3d-models-sample
CGAxis 3D Models - Free Sample (Furniture / Chairs)
A free, licensed sample of human-authored 3D models from CGAxis, a 3D content studio operating since 2008. This sample is a taster of the full CGAxis AI Data corpus (4,390 3D models + 7,794 PBR material sets) available for commercial AI-training licenses.
Every model ships as GLB and USDZ, with geometry statistics, real-world scale in centimetres, semantic tags, a natural-language caption, per-file SHA-256 and a… See the full description on the dataset page: https://huggingface.co/datasets/CGAxis/cgaxis-3d-models-sample.CGL-Dataset
Dataset Card for CGL-Dataset
Dataset Summary
CGL-Dataset is a poster layout dataset released with Composition-aware Graphic Layout GAN for Visual-Textual Presentation Designs. The paper studies layout generation for a given image, emphasizing that both global semantics and spatial image composition affect where graphic elements should be placed. The original dataset contains 60,548 advertising posters with annotated layout information.
Supported… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CGL-Dataset.CGL-Dataset-v2
Dataset Card for CGL-Dataset v2
Dataset Summary
CGL-Dataset v2 is an advertising-poster layout dataset released with Relation-Aware Diffusion Model for Controllable Poster Layout Generation. The paper argues that poster layouts should account for both visual-textual relationships and geometry relationships between elements. This version extends CGL-Dataset with richer element annotations, text annotations, and text features for controllable poster layout… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CGL-Dataset-v2.see-world-1-CGDarcticDEM-32m
ArcticDEM 32 m Mosaic
Dataset Summary
A publicly-accessible digital elevation mosaic covering the Arctic region, derived from the Polar Geospatial Center’s ArcticDEM project, at 32 m horizontal resolution.
📘 Dataset Description
SourceProvided by the University of Minnesota’s Polar Geospatial Center (PGC), part of the ArcticDEM initiative funded by NSF & NGA (website).
Spatial CoveragePan-Arctic: regions north of ~60° N (Greenland, Canada, Russia, Alaska… See the full description on the dataset page: https://huggingface.co/datasets/cgeorgiaw/arcticDEM-32m.mm_CGDSSS-GS
🕯️ Light-Stage OLAT Subsurface-Scattering Dataset
Companion data for the paper "Subsurface Scattering for 3D Gaussian Splatting"
This README documents only the dataset.A separate repo covers the training / rendering code: https://github.com/cgtuebingen/SSS-GS
Overview
Subsurface scattering (SSS) gives translucent materials (wax, soap, jade, skin) their distinctive soft glow. Our paper introduces SSS-GS, the first 3D Gaussian-Splatting framework that jointly… See the full description on the dataset page: https://huggingface.co/datasets/CGTuebingen/SSS-GS.cgaxis-pbr-materials-sample
CGAxis PBR Materials - Free Sample (Brick, Stone & Concrete)
A free, licensed sample of human-authored PBR material sets from CGAxis, a 3D content studio operating since 2008. This sample is a taster of the full CGAxis AI Data corpus (7,794 PBR material sets + 4,390 3D models) available for commercial AI-training licenses.
Every material ships as a complete, consistently named set of maps with explicit color space, bit depth and real-world scale, plus per-set metadata.json with… See the full description on the dataset page: https://huggingface.co/datasets/CGAxis/cgaxis-pbr-materials-sample.cgqa
Dataset Card for CGQA dataset
This is the CGQA dataset from the Learning Graph Embeddings for Compositional Zero-shot Learning paper.
Citation
If you use this dataset, please cite the following papers:
@inproceedings{naeem2021learning,
title={Learning graph embeddings for compositional zero-shot learning},
author={Naeem, Muhammad Ferjad and Xian, Yongqin and Tombari, Federico and Akata, Zeynep},
booktitle={Proceedings of the IEEE/CVF conference on computer… See the full description on the dataset page: https://huggingface.co/datasets/nihalnayak/cgqa.continuous-cg-full-grid
Continuous Condition-Severity nuImages Corruption Grid
This public, gated dataset contains the frozen image and metadata release used
for continuous condition/severity adapter experiments with PicoDet-M. Access is
automatically granted after a signed-in user acknowledges the terms above.
Scope
Clean nuImages source images used by the frozen training and development split.
Six derived corruption families: ColorQuant, Fog, LowLight, MotionBlur, Snow,
and… See the full description on the dataset page: https://huggingface.co/datasets/apdoa/continuous-cg-full-grid.USB
Dataset Card for USB-SafeBench
This dataset is for paper USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
You can visit our project for details at USB-SafeBench.
CGQA_and_COBJThe Official Dataset for CGQA, COBJ in NeurIPS2023 Paper "Does Continual Learning Meet Compositionality? New Benchmarks and An Evaluation Framework".Official Github Repository: Click Here
assetsCGL-Dataset
Dataset Card for CGL-Dataset
Dataset Summary
The CGL-Dataset is a dataset used for the task of automatic graphic layout design for advertising posters. It contains 61,548 samples and is provided by Alibaba Group.
Supported Tasks and Leaderboards
The task is to generate high-quality graphic layouts for advertising posters based on clean product images and their visual contents. The training set and validation set are collections of 60,548 e-commerce… See the full description on the dataset page: https://huggingface.co/datasets/qq3374129162/CGL-Dataset.cgcgschulz_bank_biotopes
Summary
Image classification dataset with biotope labels extracted from Meyer et al., 2022.
Images were extracted from six remotely operated vehicle (ROV) dives during two SponGES cruises conducted in the summers of 2017 and 2018. The ROV dives were performed by Ægir 6000 and the cruises were conducted on the RV G.O. Sars. The ROV dives traversed across various regions on the seamount, from the base of the seamount at 2700 m depth to the summit at 580 m depth. 600 images were… See the full description on the dataset page: https://huggingface.co/datasets/CGame1/schulz_bank_biotopes.twitter-sonidomix2-2026.02.23-2025811037824696440-cGDntqq2G2QTV4KO-part1interior-cgiThis new dataset contains CG interior images representing interior of houses in 5 classes, with 1000 images per class.OlympiadBenchCGRPO-20K
Dataset Name (请替换成你的数据集名称)
Dataset Description
这是一个多模态推理数据集,包含以下字段:
id:样本唯一标识(string)
problem:问题文本,通常来自复杂推理场景
data_source:原始数据来源,例如 ThinkLite-VL-Hard、Vision-SR1 等
images:图像,HuggingFace 会自动处理 {'bytes': ...} 格式
answer:标准答案
acc:模型或人工的正确率标注(0~1 float32)
你可以用于 VQA、多模态推理、模型评估等任务。
Total Records: (请填入)
Sources: ThinkLite-VL-Hard, Vision-SR1, (or others)
Tasks: Visual QA, Multimodal Reasoning, Evaluation
Dataset Structure
Data Fields
Field… See the full description on the dataset page: https://huggingface.co/datasets/KAICLIFE/CGRPO-20K.phoenix-poj104-cg-qapoj-cg-phoenix-qamm_mimic_cgdSSC-CGL-2023-hi-multimodalttbcloth2cghn
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AmeerUlAman/cghn.DyS-CGL-Datasetmimic_cgd_cleaned
mimic_cgd_cleaned
The mimic_cgd__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
51,519
QA turns
267,436
answers rewritten by the cleaning pass
64,824
QA created by the cleaning pass (new_qa)
201,168 (75.2%)
shards
10
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds wrong but… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/mimic_cgd_cleaned.phoenix-poj104-cg-qa-expandedcgebench
CGEBench
This dataset contains VLM-filtered SAM3 prompt generalization hierarchies.
Rows are accepted hierarchies from avrecum/sam3_generalizations, filtered with
google/gemma-4-31B-it through vLLM.
Each row contains:
image: embedded source image bytes.
prompt_hierarchy: prompt strings for prompt indices 0..4.
prompt_levels: per-level SAM masks in compressed COCO RLE form, mask scores,
union masks, union areas, and containment metrics.
containment_metrics: compact per-level… See the full description on the dataset page: https://huggingface.co/datasets/avrecum/cgebench.
