datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zendo-synthetic-data
Zendo Synthetic Visual Reasoning Dataset
Synthetic Zendo-style scenes with associated rules and per-scene tensor
representations. Each scene either follows ("positive", label=1) or violates
("negative", label=0) a rule that is given in natural language and as a Prolog
query.
Splits
split
scenes
train
56475
test
3344
rules total
3439
Layout
images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.synthetic_us_passports_easy
Dataset Card for synthetic_us_passports
This is a FiftyOne dataset with 9750 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/synthetic_us_passports_easy")
# Launch the App
session = fo.launch_app(dataset)
Based on the… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/synthetic_us_passports_easy.synthetic-analog-gauges
Synthetic Analog Gauges Dataset v2.0
Version: v2.0Images: 10,000 synthetic industrial gauge imagesSplits: train=8,000, val=1,000, test=1,000Resolution: 640x640 JPEGAnnotations: COCO detection/regression + semantic segmentation PNG masks + per-image metadata
What Is New In v2.0
This release is the final photorealistic industrial-background version of the VKR synthetic analog gauges dataset.
Key additions compared with the older public release:
Real generated industrial… See the full description on the dataset page: https://huggingface.co/datasets/Mileeena/synthetic-analog-gauges.frontier-synthetic-images-2026
Frontier Synthetic Images — Deduplicated Research Corpus
This is a training-only corpus of 40,290 exact-deduplicated AI-generated images from recent and frontier generators. It normalizes four provenance-pinned sources into one row-per-image schema for image-forensics research. It is not an evaluation benchmark and should not be used to report detector accuracy after training on it.
Sources and licensing
Qwen/Qwen-Image-Bench at… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/frontier-synthetic-images-2026.ddpm-derm-synthetic-gallery
Versioned synthetic dermatofibroma gallery
This repository publishes the exact, pre-generated gallery used by the
educational HAM10000 portfolio demo. It is not medical data for diagnosis or
treatment and must not be presented as clinically validated imagery.
Published version
The repository root must contain this directory unchanged:
epoch0100_seed0/
_READY.json
metadata.json
synthetic_df.csv
images/
For low-latency container builds, the repository also… See the full description on the dataset page: https://huggingface.co/datasets/sfczaa/ddpm-derm-synthetic-gallery.synthetic-turkish-passports
Turkish passport dataset - 5, 000 images
Dataset comprises 5,000 meticulously organized files capturing Turkish passports under highly controlled variations, making it an invaluable resource for developing robust document recognition and verification systems. It is specifically designed for training and testing models in passport authentication, biometric data extraction, and identity verification.
By leveraging this dataset containing detailed information from Turkish… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-turkish-passports.synthetic-human-expressions-poses-3d
3D Synthetic Human Poses and FACS Expressions Dataset
This is a high-fidelity synthetic dataset consisting of 10,075 pairs of 3D human character renders and detailed natural language annotations.
Dataset Structure & Generation
To ensure consistency, the dataset is generated using a single base 3D human model. The diversity of the dataset is achieved through a wide range of body poses, facial expressions, and camera angles:
Character: 1 base human model.
Camera… See the full description on the dataset page: https://huggingface.co/datasets/nadizik/synthetic-human-expressions-poses-3d.synthetic-object-relations
Synthetic Object Relations Dataset
A synthetic image dataset generated with Flux Schnell featuring clean object-relation prompts designed for training spatial reasoning in vision and diffusion models.
Dataset Description
This dataset contains images generated from structured prompts describing spatial relationships between objects. Unlike typical caption datasets that use free-form text, our prompts follow consistent patterns that explicitly encode:
Object identities… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-object-relations.synthetic-greeting-cards
Synthetic Greeting Cards Dataset (Synth-GCD)
A modern, fully open synthetic alternative to the proprietary Greeting Cards Dataset (GCD) described in the paper"Weakly Supervised Annotations for Multi-modal Greeting Cards Dataset".
This dataset contains high-quality AI-generated greeting card illustrations with corresponding short messages, designed for research in multimodal classification, retrieval, and generation.
Dataset Summary
Property
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/gauravs101/synthetic-greeting-cards.synthetic-characters
Synthetic Characters Dataset
A synthetic image dataset generated with Flux Schnell featuring structured character prompts designed for training character generation, fashion understanding, and portrait synthesis models.
Recommended Filters:
Age
People Count
Hair color
Camera angle
Anime/Realistic
Nudity/Clothed
I'll likely recaption everything with a list of classifications attached to them for easy filtering later.
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-characters.bilingual-ocr-ru-en-synthetic
Bilingual OCR RU-EN Synthetic Dataset
This synthetic dataset is designed for bilingual text recognition (OCR) and script classification tasks (cyrillic / latin) at the word and short-line level.
Why are numbers, mathematical symbols, and the Greek alphabet included in the generation?
When creating synthetic OCR datasets, including an expanded set of characters (digits, mathematical signs, and Greek letters) is a deliberate step aimed at two main goals:… See the full description on the dataset page: https://huggingface.co/datasets/tehnik-tehnolog/bilingual-ocr-ru-en-synthetic.dubai-taxi-advertising-synthetic
Dubai Taxi Advertising Compliance (Synthetic)
63 synthetic photorealistic images of Dubai RTA taxis carrying advertising, built
to evaluate whether vision-language models can judge out-of-home (OOH)
advertising compliance rules from a single photograph.
Each image is generated to be an unambiguous pass or fail against one
specific rule from a Dubai taxi advertising technical checklist. The dataset is
an evaluation set — it is small, adversarially balanced, and deliberately… See the full description on the dataset page: https://huggingface.co/datasets/AbdullahKhanSherwani/dubai-taxi-advertising-synthetic.synthetic-infrared-maritime-vessel-dataset-flux2klein
Synthetic Infrared Maritime Vessel Dataset (FLUX2-Klein)
RGB maritime vessel images translated to synthetic infrared via a DreamBooth-finetuned FLUX.2-Klein-4B.
Layout
synthetic-infrared-maritime-vessel-dataset-flux2klein/
├── in-distribution/
│ ├── train/{C00,C02,...}/*.jpg
│ ├── val/{C00,C02,...}/*.jpg
│ ├── test/{C00,C02,...}/*.jpg
│ ├── labels.txt
│ └── selected-metadata-{train,val,test}.json
└── out-of-distribution/
├── val/{C01,C03,C06… See the full description on the dataset page: https://huggingface.co/datasets/hanchong/synthetic-infrared-maritime-vessel-dataset-flux2klein.clide_synthetic_datasets
CLIDE Synthetic Image Datasets
📄 Paper • 💻 Code • 🌐 Webpage • 🎥 Video
A collection of synthetic images generated by modern text-to-image models, organized by domain and generator.
The dataset is designed to support analysis and evaluation of generated-image detection methods under domain and generator shifts.
🗂️ Dataset Structure
The dataset contains two visual domains:
💥🚗 Damaged Cars
Synthetic images of damaged cars generated by multiple… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/clide_synthetic_datasets.PRLx-GAN-synthetic-rim
PRLx-GAN
Repository for Synthetic Generation and Latent Projection Denoising of Rim Lesions in Multiple Sclerosis published in Synthetic Data at CVPR 2025.
Summary
Paramagnetic rim lesions (PRLs) are a rare but highly prognostic lesion subtype in multiple sclerosis, visible only on susceptibility ($\chi$) contrasts. This work presents a generative framework to:
Synthesize new rim lesion maps that address class imbalance in training data
Enable a novel denoising… See the full description on the dataset page: https://huggingface.co/datasets/agr78/PRLx-GAN-synthetic-rim.Bolt_Screw_Synthetic_Dataset
Bolt and Screw Synthetic Dataset
This dataset contains synthetic images of bolts and screws generated using Houdini.It is designed for simple computer vision experiments, especially binary image classification.
Dataset Details
Classes: 2
bolt
screw
Images per class: 500
Total images: 1,000
Image size: 512 × 512 pixels
Data type: Synthetic rendered images
Generation tool: Houdini
License: MIT
Intended Use
This dataset can be used for:
Training a small image… See the full description on the dataset page: https://huggingface.co/datasets/0707ryan/Bolt_Screw_Synthetic_Dataset.synthetic-infrared-maritime-vessel-dataset-flux2klein-degraded
Degraded Infrared Maritime Vessel Dataset
Real-ESRGAN's synthetic image degradation pipeline applied to the FLUX2-Klein synthetic infrared dataset.
Layout
synthetic-infrared-maritime-vessel-dataset-flux2klein-degraded/
├── in-distribution/
│ ├── train/{C00,C02,...}/*.jpg
│ ├── val/{C00,C02,...}/*.jpg
│ ├── test/{C00,C02,...}/*.jpg
│ ├── labels.txt
│ └── selected-metadata-{train,val,test}.json
└── out-of-distribution/
├── val/{C01,C03,C06,C16}/*.jpg… See the full description on the dataset page: https://huggingface.co/datasets/hanchong/synthetic-infrared-maritime-vessel-dataset-flux2klein-degraded.paleo-hebrew-seals-synthetic
PaleoHebrew-Seals Synthetic Corpus
This repository hosts the synthetic corpus part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions.
Why this dataset is needed
Annotated real Paleo-Hebrew seal photographs are scarce. The synthetic corpus is designed to provide large-scale supervision for training and augmentation while preserving explicit structure at the character level.
Overview
The corpus contains… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-synthetic.synthetic-chest-xray-pneumonia
Synthetic Chest X-Ray Pneumonia Dataset
Dataset Description
This dataset contains synthetic chest X-ray images generated using Stable Diffusion 2.1
fine-tuned with DreamBooth on the hf-vision/chest-xray-pneumonia dataset.
Purpose
Created for a science fair project investigating whether synthetic medical images generated
by diffusion models can improve pneumonia classifier accuracy.
Research Question
Can synthetic chest X-ray images generated by a… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/synthetic-chest-xray-pneumonia.webXOS-blackhole-synthetic
webXOS_galaxy_synthetic v1.0
_______ ___ _______ _______ ___ _ __ __ _______ ___ _______
| _ || | | _ || || | | | | | | || || | | |
| |_| || | | |_| || || |_| | | |_| || _ || | | ___|
| || | | || || _| | || | | || | | |___
| _ | | |___ | || _|| |_ | || |_| || |___ | ___|
| |_| || || _… See the full description on the dataset page: https://huggingface.co/datasets/webxos/webXOS-blackhole-synthetic.synthetic-watch-faces-dataset
Synthetic Watch Faces Dataset
A synthetic dataset of analog watch faces displaying various times for training vision models in time recognition tasks.
Dataset Description
This dataset consists of randomly generated analog watch faces showing different times. Each image contains a watch with hour and minute hands positioned to display a specific time. The dataset is designed to help train and evaluate computer vision models and Vision-Language Models (VLMs) for time… See the full description on the dataset page: https://huggingface.co/datasets/elischwartz/synthetic-watch-faces-dataset.synthetic-concrete-surface-normal
Synthetic Concrete Surface – Normal Images for Anomaly Detection
Dataset Description
A dataset of 747 synthetic images of defect-free concrete surfaces, generated for training one-class anomaly detection models such as PatchCore. All images are 1024×1024 PNG files depicting intact concrete with no cracks, spalling, or structural damage.
MVTec AD, the de facto benchmark for visual anomaly detection, covers 15 object and texture categories but does not include concrete.… See the full description on the dataset page: https://huggingface.co/datasets/wolfsynth/synthetic-concrete-surface-normal.synthetic_wildlife_health
Synthetic Wildlife Health: Camera Trap Imagery for Alopecia and Body Condition Screening
Dataset Summary
This dataset contains 553 synthetic camera trap images depicting alopecia (hair loss consistent with mange) and body condition deterioration in North American wildlife, along with paired visual question-answering annotations for health assessment tasks.
All images are AI-generated edits of real camera trap photographs sourced from iWildCam 2022. The generative pipeline… See the full description on the dataset page: https://huggingface.co/datasets/BrundageLab/synthetic_wildlife_health.camonet-synthetic
CamoNet Synthetic
A procedurally-generated military camouflage pattern dataset — 40 historical
and contemporary patterns × 200 samples each = 8,000 256×256 RGB images,
each tagged with origin, era, and visual family.
Sister project to the CamoNet model.
Where that model is trained on real photographs scraped from the web, this
dataset is fully synthetic — every image is generated from a small Python
recipe per pattern family, so the data is reproducible from a seed and
freely… See the full description on the dataset page: https://huggingface.co/datasets/Mattysmittttt/camonet-synthetic.synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/lingcarzy/synthetic-dataset-1m-dalle3-high-quality-captions.waikato_aerial_2017_synthetic_v2
Waikato Aerial Imagery 2017 Synthetic Data v2
This is a synthetic dataset generated using a sample taken from the original classification dataset residing at https://datasets.cms.waikato.ac.nz/taiao/waikato_aerial_imagery_2017/. You can find additional dataset information using the provided URL. This version (v2) has been generated using slightly altered prompts compared to v1.
Generation Params
Inference Steps: 60Images generated per prompt: 50 (1000 images per… See the full description on the dataset page: https://huggingface.co/datasets/dinushiTJ/waikato_aerial_2017_synthetic_v2.S3simulator_Plus_Synthetic_Sonar_dataset
S3Simulator+ Synthetic Sonar Dataset (Mine)
Summary
A dataset of 4,180 synthetic side-scan sonar images of the mine class (cylinderical, Truncated cone),
generated with S3Simulator+, an extended version of the S3Simulator
simulator. The dataset is described in the paper cited below. It is intended
for research on underwater image analysis when labeled real sonar data is
scarce.
The ship and plane classes are released separately, generated with
S3Simulator: see… See the full description on the dataset page: https://huggingface.co/datasets/caicv/S3simulator_Plus_Synthetic_Sonar_dataset.S3simulator_Synthetic_Sonar_dataset
S3Simulator Synthetic Sonar Dataset (Ship and Plane)
Summary
A dataset of 7,721 synthetic side-scan sonar images of two target classes,
ship and plane, generated with S3Simulator and described in the paper
cited below. It is intended as a benchmark for underwater image analysis when
labeled real sonar data is scarce.
The mine class is released separately in a companion dataset generated with
S3Simulator+.
Dataset contents
Class
Images… See the full description on the dataset page: https://huggingface.co/datasets/caicv/S3simulator_Synthetic_Sonar_dataset.myanmar-synthetic-syllable-glyphs
🇲🇲 Myanmar Synthetic Syllable Glyphs (MSSG)
The Myanmar Synthetic Syllable Glyphs (MSSG) is a massive-scale, high-fidelity synthetic image dataset containing 14,295,552 heavily augmented glyph images (128x64 pixels, grayscale) representing the structural combinatorial matrix of the Burmese script.
Developed and engineered by Khant Sint Heinn (Kalix Louis), this core foundational dataset is officially published and maintained under DatarrX (Myanmar Open Source Organization… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-synthetic-syllable-glyphs.
