datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zendo-synthetic-data
Zendo Synthetic Visual Reasoning Dataset
Synthetic Zendo-style scenes with associated rules and per-scene tensor
representations. Each scene either follows ("positive", label=1) or violates
("negative", label=0) a rule that is given in natural language and as a Prolog
query.
Splits
split
scenes
train
56475
test
3344
rules total
3439
Layout
images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.frontier-synthetic-images-2026
Frontier Synthetic Images — Deduplicated Research Corpus
This is a training-only corpus of 40,290 exact-deduplicated AI-generated images from recent and frontier generators. It normalizes four provenance-pinned sources into one row-per-image schema for image-forensics research. It is not an evaluation benchmark and should not be used to report detector accuracy after training on it.
Sources and licensing
Qwen/Qwen-Image-Bench at… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/frontier-synthetic-images-2026.bilingual-ocr-ru-en-synthetic
Bilingual OCR RU-EN Synthetic Dataset
This synthetic dataset is designed for bilingual text recognition (OCR) and script classification tasks (cyrillic / latin) at the word and short-line level.
Why are numbers, mathematical symbols, and the Greek alphabet included in the generation?
When creating synthetic OCR datasets, including an expanded set of characters (digits, mathematical signs, and Greek letters) is a deliberate step aimed at two main goals:… See the full description on the dataset page: https://huggingface.co/datasets/tehnik-tehnolog/bilingual-ocr-ru-en-synthetic.SynthCheX-75K-v2
SynthCheX-75K
SynthCheX-75K is released as a part of the CheXGenBench paper. It is a synthetic dataset generated using Sana (0.6B) [1] fine-tuned on chest radiographs. Sana (0.6B) establishes the SoTA performance on the CheXGenBench benchmark.
The dataset contains 75,649 high-quality image-text samples along with the pathological annotations.
Filtration Process for SynthCheX-75K
Generative models can lead to both high and low-fidelity generations on different subsets… See the full description on the dataset page: https://huggingface.co/datasets/raman07/SynthCheX-75K-v2.synthetic-characters
Synthetic Characters Dataset
A synthetic image dataset generated with Flux Schnell featuring structured character prompts designed for training character generation, fashion understanding, and portrait synthesis models.
Recommended Filters:
Age
People Count
Hair color
Camera angle
Anime/Realistic
Nudity/Clothed
I'll likely recaption everything with a list of classifications attached to them for easy filtering later.
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-characters.synthetic-object-relations
Synthetic Object Relations Dataset
A synthetic image dataset generated with Flux Schnell featuring clean object-relation prompts designed for training spatial reasoning in vision and diffusion models.
Dataset Description
This dataset contains images generated from structured prompts describing spatial relationships between objects. Unlike typical caption datasets that use free-form text, our prompts follow consistent patterns that explicitly encode:
Object identities… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-object-relations.africa-synth-aid-flows-brain-tumor-mri-colorized-ehr-all
Brain Tumor (MRI) Detection Colourized with EHR | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: not declared - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-brain-tumor-mri-colorized-ehr-all.synthetic-chest-xray-pneumonia
Synthetic Chest X-Ray Pneumonia Dataset
Dataset Description
This dataset contains synthetic chest X-ray images generated using Stable Diffusion 2.1
fine-tuned with DreamBooth on the hf-vision/chest-xray-pneumonia dataset.
Purpose
Created for a science fair project investigating whether synthetic medical images generated
by diffusion models can improve pneumonia classifier accuracy.
Research Question
Can synthetic chest X-ray images generated by a… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/synthetic-chest-xray-pneumonia.synthetic-watch-faces-dataset
Synthetic Watch Faces Dataset
A synthetic dataset of analog watch faces displaying various times for training vision models in time recognition tasks.
Dataset Description
This dataset consists of randomly generated analog watch faces showing different times. Each image contains a watch with hour and minute hands positioned to display a specific time. The dataset is designed to help train and evaluate computer vision models and Vision-Language Models (VLMs) for time… See the full description on the dataset page: https://huggingface.co/datasets/elischwartz/synthetic-watch-faces-dataset.paleo-hebrew-seals-synthetic
PaleoHebrew-Seals Synthetic Corpus
This repository hosts the synthetic corpus part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions.
Why this dataset is needed
Annotated real Paleo-Hebrew seal photographs are scarce. The synthetic corpus is designed to provide large-scale supervision for training and augmentation while preserving explicit structure at the character level.
Overview
The corpus contains… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-synthetic.synthetic_wildlife_health
Synthetic Wildlife Health: Camera Trap Imagery for Alopecia and Body Condition Screening
Dataset Summary
This dataset contains 553 synthetic camera trap images depicting alopecia (hair loss consistent with mange) and body condition deterioration in North American wildlife, along with paired visual question-answering annotations for health assessment tasks.
All images are AI-generated edits of real camera trap photographs sourced from iWildCam 2022. The generative pipeline… See the full description on the dataset page: https://huggingface.co/datasets/BrundageLab/synthetic_wildlife_health.waikato_aerial_2017_synthetic_v2
Waikato Aerial Imagery 2017 Synthetic Data v2
This is a synthetic dataset generated using a sample taken from the original classification dataset residing at https://datasets.cms.waikato.ac.nz/taiao/waikato_aerial_imagery_2017/. You can find additional dataset information using the provided URL. This version (v2) has been generated using slightly altered prompts compared to v1.
Generation Params
Inference Steps: 60Images generated per prompt: 50 (1000 images per… See the full description on the dataset page: https://huggingface.co/datasets/dinushiTJ/waikato_aerial_2017_synthetic_v2.camonet-synthetic
CamoNet Synthetic
A procedurally-generated military camouflage pattern dataset — 40 historical
and contemporary patterns × 200 samples each = 8,000 256×256 RGB images,
each tagged with origin, era, and visual family.
Sister project to the CamoNet model.
Where that model is trained on real photographs scraped from the web, this
dataset is fully synthetic — every image is generated from a small Python
recipe per pattern family, so the data is reproducible from a seed and
freely… See the full description on the dataset page: https://huggingface.co/datasets/Mattysmittttt/camonet-synthetic.brain-tumor-single-slice-MRI-scan-with-synthetic-ehr-africa
Dataset Card: Africa Brain Tumor Scans with Synthetic EHR (Bundled Parquet)
This dataset bundles single-slice brain MRI scans and richly structured, synthetic EHR data into a single Parquet file suitable for multimodal ML research. Each row contains an image struct (bytes + path), a source label column, and an EHR payload with both a full JSON record and convenient summary columns.
The synthetic EHRs are Africa-focused: they encode country, urban/rural, facility level, insurance… See the full description on the dataset page: https://huggingface.co/datasets/saad02/brain-tumor-single-slice-MRI-scan-with-synthetic-ehr-africa.myanmar-synthetic-syllable-glyphs
🇲🇲 Myanmar Synthetic Syllable Glyphs (MSSG)
The Myanmar Synthetic Syllable Glyphs (MSSG) is a massive-scale, high-fidelity synthetic image dataset containing 14,295,552 heavily augmented glyph images (128x64 pixels, grayscale) representing the structural combinatorial matrix of the Burmese script.
Developed and engineered by Khant Sint Heinn (Kalix Louis), this core foundational dataset is officially published and maintained under DatarrX (Myanmar Open Source Organization… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-synthetic-syllable-glyphs.synthetic-shapes-3x6x7
Synthetic Shapes 3×6×7
A fully deterministic synthetic dataset of simple geometric shapes rendered as SVG images, with precomputed CLIP (ViT-B-32) embeddings for both text and images.
Purpose
This dataset is designed for controlled experiments in representation alignment and steering vector evaluation. Because images are generated deterministically from a known combinatorial space, it provides a clean testbed where ground-truth structure is fully known.… See the full description on the dataset page: https://huggingface.co/datasets/amirali1985/synthetic-shapes-3x6x7.synthetic-digital-signature-blocks
Synthetic Digital Signature Blocks
Synthetically generated images of digital signature appearance blocks — the visual stamp that PDF readers draw on a document when it is signed electronically. Every image is procedurally generated with PIL; no real document, no real person, and no real product branding is included.
The dataset was built to train signature-presence detectors on scanned administrative forms, where handwritten signatures are only one of several valid visual… See the full description on the dataset page: https://huggingface.co/datasets/marayagomez/synthetic-digital-signature-blocks.waikato_aerial_2017_synthetic_best_cmmdwaikato_aerial_2017_synthetic_v1
Waikato Aerial Imagery 2017 Synthetic Data v1
This is a synthetic dataset generated using a sample taken from the original classification dataset residing at https://datasets.cms.waikato.ac.nz/taiao/waikato_aerial_imagery_2017/. You can find additional dataset information using the provided URL.
Generation Params
Inference Steps: 60Images generated per prompt: 50 (1000 images per class since there are 20 prompts for each class)
CMMD (CLIP Maximum Mean… See the full description on the dataset page: https://huggingface.co/datasets/dinushiTJ/waikato_aerial_2017_synthetic_v1.waikato_aerial_2017_synthetic_best_fidwaikato_aerial_2017_synthetic_v0
Waikato Aerial Imagery 2017 Synthetic Data v0
This is a synthetic dataset generated using a sample taken from the original classification dataset residing at https://datasets.cms.waikato.ac.nz/taiao/waikato_aerial_imagery_2017/. You can find additional dataset information using the provided URL.
Generation Params
Inference Steps: 30Images generated per prompt: 50 (1000 images per class since there are 20 prompts for each class)
CMMD (CLIP Maximum Mean… See the full description on the dataset page: https://huggingface.co/datasets/dinushiTJ/waikato_aerial_2017_synthetic_v0.car-ukraine-synth
Ukraine Synthetic Vehicle Dataset — Toyota Corolla × BMW 3 Series
Фотореалістичні синтетичні зображення седанів Toyota Corolla та BMW 3 Series в українських урбаністичних сценах, згенеровані через OpenAI gpt-image-2. Кожне зображення семплить з сітки 12 українських міст × 7 погодних умов × 5 часів доби × 6 типів камер (dashcam, CCTV, drone, smartphone, action cam, wall-mounted security).
Зображень: 150
Джерело: OpenAI gpt-image-2 (reference-conditioned edits endpoint)… See the full description on the dataset page: https://huggingface.co/datasets/arrmlet/car-ukraine-synth.heb_synth_pangoline
Dataset Card for Hebrew Synthetic Pangoline Dataset
INFO: I'm not giving access to users with 0 models/0 datasets/0 activity - sharing is both ways
Dataset Summary
The Hebrew Synthetic Pangoline Dataset is a comprehensive collection of synthetic Hebrew document images generated using a custom implementation of Pangoline, a text-to-image synthesis tool. The dataset contains high-quality synthetic Hebrew text rendered as images, along with corresponding ground truth… See the full description on the dataset page: https://huggingface.co/datasets/johnlockejrr/heb_synth_pangoline.synthetic-face-sdxl-instantid-bench
Synthetic Face Detection Benchmark — SDXL+InstantID
Version: v1.0.0 · Build date: 2026-05-16 · Rows: 26492
Evaluation benchmark for synthetic-face detection under platform-realistic
conditions. Sampled to satisfy the ISO/IEC 19795 floor of 300 samples per
demographic subgroup across a 6×2 (skin tone × gender) cell grid. Not
training data; not licensed for commercial use.
See release.json for build provenance, manifest.csv for per-row
license attestation, LICENSES.csv for the… See the full description on the dataset page: https://huggingface.co/datasets/danb21/synthetic-face-sdxl-instantid-bench.yid_synth_pangolineINFO: I'm not giving access to users with 0 models/0 datasets/0 activity - sharing is both ways
Dataset Summary
The Yiddish Synthetic Pangoline Dataset is a comprehensive collection of synthetic Yiddish document images generated using a custom implementation of Pangoline, a text-to-image synthesis tool. The dataset contains high-quality synthetic Yiddish text rendered as images, along with corresponding ground truth text and ALTO-XML layout annotations. This dataset is designed… See the full description on the dataset page: https://huggingface.co/datasets/johnlockejrr/yid_synth_pangoline.
