datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zendo-synthetic-data
Zendo Synthetic Visual Reasoning Dataset
Synthetic Zendo-style scenes with associated rules and per-scene tensor
representations. Each scene either follows ("positive", label=1) or violates
("negative", label=0) a rule that is given in natural language and as a Prolog
query.
Splits
split
scenes
train
56475
test
3344
rules total
3439
Layout
images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.synthetic_us_passports_easy
Dataset Card for synthetic_us_passports
This is a FiftyOne dataset with 9750 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/synthetic_us_passports_easy")
# Launch the App
session = fo.launch_app(dataset)
Based on the… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/synthetic_us_passports_easy.ddpm-derm-synthetic-gallery
Versioned synthetic dermatofibroma gallery
This repository publishes the exact, pre-generated gallery used by the
educational HAM10000 portfolio demo. It is not medical data for diagnosis or
treatment and must not be presented as clinically validated imagery.
Published version
The repository root must contain this directory unchanged:
epoch0100_seed0/
_READY.json
metadata.json
synthetic_df.csv
images/
For low-latency container builds, the repository also… See the full description on the dataset page: https://huggingface.co/datasets/sfczaa/ddpm-derm-synthetic-gallery.frontier-synthetic-images-2026
Frontier Synthetic Images — Deduplicated Research Corpus
This is a training-only corpus of 40,290 exact-deduplicated AI-generated images from recent and frontier generators. It normalizes four provenance-pinned sources into one row-per-image schema for image-forensics research. It is not an evaluation benchmark and should not be used to report detector accuracy after training on it.
Sources and licensing
Qwen/Qwen-Image-Bench at… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/frontier-synthetic-images-2026.mtg_synthetic_large_dataset
Magic The Gatering Synthetic Large Image Dataset.
Synthetic images of MTG cards with the following features:
Photorealistic rendering using BlenderProc2 and HDRI environments
Precise card geometry with rounded corners
Random transformations for data augmentation
Segmentation masks for semantic segmentation training
173 GB of data, 718508 images for train, 307932 images for testing
for uncompress in ubuntu:
sudo apt install p7zip-full
7z x mtg_synthetic_large_dataset.7z.001
Source… See the full description on the dataset page: https://huggingface.co/datasets/dhvazquez/mtg_synthetic_large_dataset.synthetic-analog-gauges
Synthetic Analog Gauges Dataset v2.0
Version: v2.0Images: 10,000 synthetic industrial gauge imagesSplits: train=8,000, val=1,000, test=1,000Resolution: 640x640 JPEGAnnotations: COCO detection/regression + semantic segmentation PNG masks + per-image metadata
What Is New In v2.0
This release is the final photorealistic industrial-background version of the VKR synthetic analog gauges dataset.
Key additions compared with the older public release:
Real generated industrial… See the full description on the dataset page: https://huggingface.co/datasets/Mileeena/synthetic-analog-gauges.synthetic-turkish-passports
Turkish passport dataset - 5, 000 images
Dataset comprises 5,000 meticulously organized files capturing Turkish passports under highly controlled variations, making it an invaluable resource for developing robust document recognition and verification systems. It is specifically designed for training and testing models in passport authentication, biometric data extraction, and identity verification.
By leveraging this dataset containing detailed information from Turkish… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-turkish-passports.synthetic-human-expressions-poses-3d
3D Synthetic Human Poses and FACS Expressions Dataset
This is a high-fidelity synthetic dataset consisting of 10,075 pairs of 3D human character renders and detailed natural language annotations.
Dataset Structure & Generation
To ensure consistency, the dataset is generated using a single base 3D human model. The diversity of the dataset is achieved through a wide range of body poses, facial expressions, and camera angles:
Character: 1 base human model.
Camera… See the full description on the dataset page: https://huggingface.co/datasets/nadizik/synthetic-human-expressions-poses-3d.synthetic-australian-medical-documents-sample
Synthetic Australian Medical Documents - Sample
A 50-document free sample of a 5,000-document library of synthetic Australian medical PDFs. PHI-free. Modelled on Australian healthcare documentation. Pre-labelled with structured ground truth and pixel-precise bounding boxes. Released under CC-BY-NC 4.0 for evaluation and non-commercial research.
See Pricing & licensing below.
What's in this sample
Field
Value
Documents
50
Document types
29 (of 45 in full… See the full description on the dataset page: https://huggingface.co/datasets/RootCauseAnalytics/synthetic-australian-medical-documents-sample.bilingual-ocr-ru-en-synthetic
Bilingual OCR RU-EN Synthetic Dataset
This synthetic dataset is designed for bilingual text recognition (OCR) and script classification tasks (cyrillic / latin) at the word and short-line level.
Why are numbers, mathematical symbols, and the Greek alphabet included in the generation?
When creating synthetic OCR datasets, including an expanded set of characters (digits, mathematical signs, and Greek letters) is a deliberate step aimed at two main goals:… See the full description on the dataset page: https://huggingface.co/datasets/tehnik-tehnolog/bilingual-ocr-ru-en-synthetic.SynthCheX-75K-v2
SynthCheX-75K
SynthCheX-75K is released as a part of the CheXGenBench paper. It is a synthetic dataset generated using Sana (0.6B) [1] fine-tuned on chest radiographs. Sana (0.6B) establishes the SoTA performance on the CheXGenBench benchmark.
The dataset contains 75,649 high-quality image-text samples along with the pathological annotations.
Filtration Process for SynthCheX-75K
Generative models can lead to both high and low-fidelity generations on different subsets… See the full description on the dataset page: https://huggingface.co/datasets/raman07/SynthCheX-75K-v2.synthetic-characters
Synthetic Characters Dataset
A synthetic image dataset generated with Flux Schnell featuring structured character prompts designed for training character generation, fashion understanding, and portrait synthesis models.
Recommended Filters:
Age
People Count
Hair color
Camera angle
Anime/Realistic
Nudity/Clothed
I'll likely recaption everything with a list of classifications attached to them for easy filtering later.
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-characters.synthetic-object-relations
Synthetic Object Relations Dataset
A synthetic image dataset generated with Flux Schnell featuring clean object-relation prompts designed for training spatial reasoning in vision and diffusion models.
Dataset Description
This dataset contains images generated from structured prompts describing spatial relationships between objects. Unlike typical caption datasets that use free-form text, our prompts follow consistent patterns that explicitly encode:
Object identities… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-object-relations.open-set-synth-img-attributionThis dataset contains synthetic images generated by different architectures trained on different datasets. There are two types of labels: architecture and generator. The architecture labels are the architectures of the generators used to synthesize the images. The generator labels are a tuple of (architecture, training_dataset) used to synthesize the images. The dataset is split into train, validation, and test sets. For more information, see the [BMVC 2023 paper](https://papers.bmvc2023.org/0659.pdf).synthetic-greeting-cards
Synthetic Greeting Cards Dataset (Synth-GCD)
A modern, fully open synthetic alternative to the proprietary Greeting Cards Dataset (GCD) described in the paper"Weakly Supervised Annotations for Multi-modal Greeting Cards Dataset".
This dataset contains high-quality AI-generated greeting card illustrations with corresponding short messages, designed for research in multimodal classification, retrieval, and generation.
Dataset Summary
Property
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/gauravs101/synthetic-greeting-cards.synthid-research
SynthID Research Set
The same prompts rendered by several generators. Some carry an invisible provenance watermark
and one does not, so the set is a matched comparison for anyone working on AI-image provenance
detection.
We are publishing it because generating it is not free, and because a matched set — same prompt,
several renderers, one of them a clean control — is more useful than unrelated collections.
This dataset grows. Directories are added as new models ship and existing… See the full description on the dataset page: https://huggingface.co/datasets/bigdatamark/synthid-research.synthetic-infrared-maritime-vessel-dataset-flux2klein
Synthetic Infrared Maritime Vessel Dataset (FLUX2-Klein)
RGB maritime vessel images translated to synthetic infrared via a DreamBooth-finetuned FLUX.2-Klein-4B.
Layout
synthetic-infrared-maritime-vessel-dataset-flux2klein/
├── in-distribution/
│ ├── train/{C00,C02,...}/*.jpg
│ ├── val/{C00,C02,...}/*.jpg
│ ├── test/{C00,C02,...}/*.jpg
│ ├── labels.txt
│ └── selected-metadata-{train,val,test}.json
└── out-of-distribution/
├── val/{C01,C03,C06… See the full description on the dataset page: https://huggingface.co/datasets/hanchong/synthetic-infrared-maritime-vessel-dataset-flux2klein.africa-synth-tuberculosis-chest-ctscan-african-ehr-all
Chest CT Scans + Synthetic African EHR (Lung Cancer) | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: imagefolder - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-tuberculosis-chest-ctscan-african-ehr-all.PRLx-GAN-synthetic-rim
PRLx-GAN
Repository for Synthetic Generation and Latent Projection Denoising of Rim Lesions in Multiple Sclerosis published in Synthetic Data at CVPR 2025.
Summary
Paramagnetic rim lesions (PRLs) are a rare but highly prognostic lesion subtype in multiple sclerosis, visible only on susceptibility ($\chi$) contrasts. This work presents a generative framework to:
Synthesize new rim lesion maps that address class imbalance in training data
Enable a novel denoising… See the full description on the dataset page: https://huggingface.co/datasets/agr78/PRLx-GAN-synthetic-rim.Crop-Disease-Image-Eval-Synthetic
Crop, Category, Disease and Pest Test Set
11,057 smallholder-farmer photographs sent to FarmerChat from Ethiopia, India, Kenya and Nigeria, each
labelled with the crop, whether the problem is a disease or a pest, and which one. This is the held-out
test split of a four-head classification benchmark, restricted to the rows whose labels came from an
independent model council rather than from the production vendor.
Why 11,057 and not 16,275
The full held-out split is… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Crop-Disease-Image-Eval-Synthetic.africa-synth-aid-flows-brain-tumor-mri-colorized-ehr-all
Brain Tumor (MRI) Detection Colourized with EHR | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: not declared - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-brain-tumor-mri-colorized-ehr-all.synthetic-chest-xray-pneumonia
Synthetic Chest X-Ray Pneumonia Dataset
Dataset Description
This dataset contains synthetic chest X-ray images generated using Stable Diffusion 2.1
fine-tuned with DreamBooth on the hf-vision/chest-xray-pneumonia dataset.
Purpose
Created for a science fair project investigating whether synthetic medical images generated
by diffusion models can improve pneumonia classifier accuracy.
Research Question
Can synthetic chest X-ray images generated by a… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/synthetic-chest-xray-pneumonia.Bolt_Screw_Synthetic_Dataset
Bolt and Screw Synthetic Dataset
This dataset contains synthetic images of bolts and screws generated using Houdini.It is designed for simple computer vision experiments, especially binary image classification.
Dataset Details
Classes: 2
bolt
screw
Images per class: 500
Total images: 1,000
Image size: 512 × 512 pixels
Data type: Synthetic rendered images
Generation tool: Houdini
License: MIT
Intended Use
This dataset can be used for:
Training a… See the full description on the dataset page: https://huggingface.co/datasets/0707ryan/Bolt_Screw_Synthetic_Dataset.synthetic-infrared-maritime-vessel-dataset-flux2klein-degraded
Degraded Infrared Maritime Vessel Dataset
Real-ESRGAN's synthetic image degradation pipeline applied to the FLUX2-Klein synthetic infrared dataset.
Layout
synthetic-infrared-maritime-vessel-dataset-flux2klein-degraded/
├── in-distribution/
│ ├── train/{C00,C02,...}/*.jpg
│ ├── val/{C00,C02,...}/*.jpg
│ ├── test/{C00,C02,...}/*.jpg
│ ├── labels.txt
│ └── selected-metadata-{train,val,test}.json
└── out-of-distribution/
├── val/{C01,C03,C06,C16}/*.jpg… See the full description on the dataset page: https://huggingface.co/datasets/hanchong/synthetic-infrared-maritime-vessel-dataset-flux2klein-degraded.synthetic-human-portrait-attributes
AI-generated potraits of people
It's not perfect, but could be useful for training for example hairstyle and color classification!
Includes 6 categories of labels:
6 ages: 'adult', 'elderly', 'mature', 'teenage', 'young', 'young_adult'
2 sexes: 'man', 'woman'
4 hair lengths: 'long', 'short', 'buzzcut', 'bald'
4 hair shapes: 'curly', 'straight', 'wavy', 'None' (for bald)
9 hair colors: 'black', 'blonde', 'platinum_blonde', 'brunette', 'ginger', 'multicolored', 'None' (for bald)
5… See the full description on the dataset page: https://huggingface.co/datasets/Zhincore/synthetic-human-portrait-attributes.synthetic-watch-faces-dataset
Synthetic Watch Faces Dataset
A synthetic dataset of analog watch faces displaying various times for training vision models in time recognition tasks.
Dataset Description
This dataset consists of randomly generated analog watch faces showing different times. Each image contains a watch with hour and minute hands positioned to display a specific time. The dataset is designed to help train and evaluate computer vision models and Vision-Language Models (VLMs) for time… See the full description on the dataset page: https://huggingface.co/datasets/elischwartz/synthetic-watch-faces-dataset.dubai-taxi-advertising-synthetic
Dubai Taxi Advertising Compliance (Synthetic)
63 synthetic photorealistic images of Dubai RTA taxis carrying advertising, built
to evaluate whether vision-language models can judge out-of-home (OOH)
advertising compliance rules from a single photograph.
Each image is generated to be an unambiguous pass or fail against one
specific rule from a Dubai taxi advertising technical checklist. The dataset is
an evaluation set — it is small, adversarially balanced, and deliberately… See the full description on the dataset page: https://huggingface.co/datasets/AbdullahKhanSherwani/dubai-taxi-advertising-synthetic.Synth-Text-Eng-512x128
Synthetic Text Images (English)
A synthetic dataset of rendered text images with rich per-sample
annotations: the text itself, its rendering attributes, background
description, applied post-processing, and a natural-language caption.
Each image is generated by compositing English text over a procedurally
generated background with random font, color, position, rotation, blur,
brightness and noise. All samples are accompanied by a structured
metadata.csv and a ready-to-use… See the full description on the dataset page: https://huggingface.co/datasets/Nininkkka/Synth-Text-Eng-512x128.SynthBench
SynthBench
Benchmark dataset for evaluating synthetic data generation methods for visual classification. Contains real iPhone photos, FLUX text-to-image generated images, and programmatically augmented synthetic images across 6 object classes.
Classes
mouse, pen, phone, laptop, water bottle, Rubik's cube
Dataset Structure
data/
├── raw/ # 308 raw iPhone photos (HEIC converted to JPEG)
├── real/ # 309 preprocessed images… See the full description on the dataset page: https://huggingface.co/datasets/LakshC/SynthBench.
