ud-synthetic/printed-usa-passports
Introduction The Synthetic Printed USA Passports Dataset contains 9,600 AI-generated passport images designed for training OCR and computer vision models on identity documents. The dataset includes varied angles, lighting conditions, backgrounds, and distances, with structured metadata covering gender, age group, resolution, and more. All images are synthetically generated — no real personal data or biometric records are involved — making it a privacy-compliant solution for… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/printed-usa-passports.
Introduction
The Synthetic Printed USA Passports Dataset contains 9,600 AI-generated passport images designed for training OCR and computer vision models on identity documents. The dataset includes varied angles, lighting conditions, backgrounds, and distances, with structured metadata covering gender, age group, resolution, and more. All images are synthetically generated — no real personal data or biometric records are involved — making it a privacy-compliant solution for building identity verification and fraud detection systems.- [Get the data](https://unidata.pro/datasets/synthetic-printed-usa-passports-dataset/?utm_source=huggingface-synthetic&utm_medium=referral&utm_campaign=synthetic-printed-usa-passports)
Dataset general info
The dataset consists of 9,600 synthetic passport images organized into sets of 96 samples, each covering every combination of capture angle, lighting, background, and distance.
Use cases
Training document verification systems Banks, fintech platforms, and border control teams can use this dataset to train models that verify printed identification documents across diverse real-world conditions. The range of passport photos with varied lighting and angles helps verification systems handle edge cases that cause failures in production.
Building fraud detection pipelines Security and compliance teams building fraud detection or attack detection models can use these synthetic datasets to generate realistic negative and positive training samples without sourcing real documents. Because no personal data is involved, teams can scale training data freely while staying compliant.
FAQ
Is this real-world or synthetic data? All 9,600 images are AI-generated and contain no biometric data or personal information tied to real individuals.
Can I request a sample before purchasing? Yes — free samples are available so you can evaluate image quality, metadata structure, and variation coverage before committing.
How is the dataset delivered? After purchase, the full dataset is delivered within 3–10 business days via secure AWS cloud infrastructure compliant with ISO 27001 and ISO 27701.
