datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GenEvolve-Data-Bench
GenEvolve Data and Bench
This repository contains the open-source data release for GenEvolve:
Config
Directory
Records
Images
Purpose
sft
GenEvolve-Data-SFT/
9,000 trajectories
50,291 reference images
supervised cold-start trajectories
rl
GenEvolve-Data-RL/
3,175 prompts
3,175 GT images
self-evolution / RL training prompts
bench
GenEvolve-Bench/
594 prompts
594 GT images
held-out evaluation benchmarkAll metadata is provided in both JSONL and Parquet. The Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/MeiGen-AI/GenEvolve-Data-Bench.assetsGenECGGenECG is an image-based ECG dataset which has been created from the PTB-XL dataset (https://physionet.org/content/ptb-xl/1.0.3/).
The PTB-XL dataset is a signal-based ECG dataset comprising 21799 unique ECGs.
GenECG is divided into the following subsets:
-Dataset A: ECGs without imperfections (Dataset_A_ECGs_without_imperfections) - This subset includes 21799 ECG images that have been generated directly from the original PTB-XL recordings, free from any visual imperfections.
-Dataset B: ECGs… See the full description on the dataset page: https://huggingface.co/datasets/edcci/GenECG.GenImageUI-Genie-Agent-16kThis repository contains the Trajectory dataset from the paper UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based
Mobile GUI Agents.
Github: https://github.com/Euphoria16/UI-Genie
GenPoster100K
Dataset Card for GenPoster100K
Dataset Summary
GenPoster-100K is a large-scale dataset for content-aware graphic layout generation introduced in the SEGA paper.
The paper describes it as a high-quality poster dataset with layer-parseable source materials and rich metadata.
This repository provides a Hugging Face datasets loader implementation that reads the source release (BruceW91/GenPoster-100K) and exposes normalized examples with:
poster background image… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/GenPoster100K.stable-bias-generationsflux_generatedNameonly_generated
Just Say the Name: Online Continual Learning with Category Names Only via Data Generation
We provide the dataset used for Name-only continual learning, generated using Stable Diffusion XL, DALL.E-2, CogView2, and DeepFloyd IF models.
Disclaimer
This dataset is created solely for academic purposes. We minimized human intervention to ensure a fair comparison with the baseline methods discussed in our paper. Despite our efforts, the extensive size of the dataset prevented us… See the full description on the dataset page: https://huggingface.co/datasets/seongwon980/Nameonly_generated.GenIRGenFig1cad-gen-freecad
CAD Generation Dataset
Each row in this dataset describes one parametric CAD part. Columns:
id — row identifier (also the basename of the per-row asset files)
name — part family (e.g. flanges, spur_gear_stock)
description — natural-language description of the geometry
key_parameters — the dimensions that drive the parametric model
image — 512×512 PNG preview rendered from the FCStd
fcstd_path — relative path inside this repo to the parametric FreeCAD document (fcstd/<id>.FCStd)… See the full description on the dataset page: https://huggingface.co/datasets/gnucleus-ai/cad-gen-freecad.OpenING_Generators_Outputslibero_gen_goal_chain_train_openpiGenEmotions-RTX5070Tiphysical-ai-bench-generation
Physical AI Bench - Generation
Paper | Code
Dataset Description
The PAI-Bench is a benchmark to measure the progress of world models quantitatively.
The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.ARTO-Gen-Dataset
ARTO-KG: A Synthetic Artwork Dataset for Knowledge-Enhanced Understanding
Dataset Description
ARTO-KG is a large-scale synthetic artwork dataset that bridges visual content and structured knowledge through ontology-guided automated generation. Each artwork is annotated with comprehensive RDF knowledge graphs aligned with the ARTO ontology.
Dataset Summary
Total Artworks: 10,108 high-resolution images (1024×1024)
Object Instances: 39,878 (average… See the full description on the dataset page: https://huggingface.co/datasets/youngcan1/ARTO-Gen-Dataset.Tiny-GenImage
Tiny GenImage Dataset
📝 Dataset Description
Dataset Summary
The Tiny GenImage Dataset is a curated, scaled-down collection of images and associated metadata designed to train, validate, and benchmark models for detecting and identifying artificially generated content. The dataset contains a mix of real-world images alongside those generated by prominent AI models, including various diffusion models (like Stable Diffusion 1.4/1.5, GLIDE, Midjourney, ADM, VQDM… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/Tiny-GenImage.generic_character_skins
Generic Character Skins Dataset
Summary
This comprehensive dataset provides an extensive collection of character images sourced from Zerochan across multiple popular anime, game, and manga franchises. The dataset contains meticulously organized character artwork spanning diverse genres including gacha games, idol franchises, fantasy series, and action RPGs. With over 2,000 character folders and thousands of high-quality images, this repository serves as a valuable… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/generic_character_skins.gently-perception-benchmark
Gently Perception Agent Benchmark
Light-sheet microscopy volumes of C. elegans embryo development, intended
for evaluating vision-based perception agents on embryo stage classification.
The dataset has two tiers:
Annotated benchmark set (embryo_1–embryo_8) — human ground-truth
stage transitions. Use this for evaluation.
Unannotated corpus (embryo_9–embryo_105) — 97 additional real embryo
timelapses with no human labels, provided for developing and stress-
testing perception… See the full description on the dataset page: https://huggingface.co/datasets/gently-project/gently-perception-benchmark.MLLM-Generated-Image-Detection-Dataset
MLLM-Generated Image Dataset
This dataset contains real and AI-generated image samples organized for binary MLLM-generated image detection.
Paper | Code
Dataset Summary
We construct an MLLM-generated image detection benchmark from GPT Image2 and Nano Banana2. This benchmark covers texture-dominated, structure-dominated, and hybrid-dominated. It is designed to evaluate detector performance under the new challenges introduced by large-scale image generation models.… See the full description on the dataset page: https://huggingface.co/datasets/zr-zhang/MLLM-Generated-Image-Detection-Dataset.Face-Gender-Swap
Dataset Card for "Face-Gender-Swap"
More Information needed
generated-imagesgenerated-passport-faces-aditya-second-halfReport-Generator-DataGMAI-Reasoning10K
GMAI-Reasoning10K
Medical Reasoning dataset used in GMAI-VL-R1
Data description
GMAI-Reasoning10K is a high-quality medical image reasoning dataset containing 10,000 carefully selected samples. The data was collected from 95 medical datasets from reliable sources such as Kaggle, GrandChallenge, and Open-Release, covering 12 imaging modalities including X-ray, CT, and MRI.
Data preprocessing followed the standardization methods from SAMed-20M: 3D data (CT/MRI) had… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-Reasoning10K.genhome3d-1280
GenHome3D-1280
1,280 validated household and spatial-design assets in USDZ format, organized
across 64 categories.
Explore the visual catalog ·
Browse the GitHub repository ·
Download the versioned release ·
Read the generation method
Dataset summary
Assets
1,280
Categories
64
Assets per category
20
Runtime format
USDZ
Units
Meters
Asset license
CC BY 4.0
Technical validation
1,280/1,280 pass
Package validation
1… See the full description on the dataset page: https://huggingface.co/datasets/linxy97/genhome3d-1280.SDXL-Generated-Stanford-Dogs
Dataset Card for Generated Dogs
10+ images for each class in the Stanford Dogs dataset, but all generated with SDXL. Images were filtered for CLIP score, and cartoonish images were removed.
This is a FiftyOne dataset with 1305 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/SDXL-Generated-Stanford-Dogs.GenSC-6G
GenSC-6G - Scalable Semantic Communication Framework and Dataset
This repository contains the first semantic communication dataset and playground, designed to be scalable, reproducible, and adaptable for a wide range of applications. The dataset and framework are tailored for semantic decoding, classification, and localization tasks in 6G applications, integrating generative AI and semantic communication. Implementation of GenSC-6G: A Prototype Testbed for Integrated Generative AI… See the full description on the dataset page: https://huggingface.co/datasets/CQILAB/GenSC-6G.
