ud-synthetic/japanese-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction The Synthetic Japan Passports Dataset brings together more than 1,000 AI-generated passport images, purpose-built for training OCR and computer vision models on identity documents. Since the data is entirely synthetic —… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/japanese-passports.
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction
The Synthetic Japan Passports Dataset brings together more than 1,000 AI-generated passport images, purpose-built for training OCR and computer vision models on identity documents. Since the data is entirely synthetic — with no genuine personal records or biometric details involved — it provides a privacy-safe option for teams developing identity verification and fraud detection solutions. - [Get the data](https://unidata.pro/datasets/synthetic-passports/?utm_source=huggingface-synthetic&utm_medium=referral&utm_campaign=synthetic-japan-passports)
The dataset can be expanded to fit your needs — additional images and metadata can be generated to your specifications, with new samples typically ready within one week.
Imagery covers both clean white backgrounds and a variety of realistic everyday scenes, including desks, walls, and other common surfaces, giving models the visual diversity needed to perform well in production environments.
The dataset spans 50+ countries (including China, South Korea, Singapore, Indonesia, Vietnam, and more). To learn more, please submit a request via [the website](https://unidata.pro/datasets/synthetic-passports/?utm_source=huggingface-synthetic&utm_medium=referral&utm_campaign=synthetic-japan-passports) to learn more.
Each image comes with comprehensive structured metadata that captures personal document fields — passport number, full name, signature, date of birth, sex, place of birth, issuing authority, nationality, document type, and machine-readable zone (MRZ) — together with technical attributes such as resolution and category.
Dataset general info
The dataset consists of 1000+ synthetic passport images, each associated with complete identity-like records and structured annotations. Additional data can be generated upon request within one week. The images cover both clean white and diverse background scenarios.
Metadata fields include:
Use cases
Training document verification systems
Banks, fintech operators, and immigration authorities can apply this dataset to develop models that authenticate identity documents under varied real-world conditions. The structured metadata enables accurate field-level extraction and validation, while the assortment of background settings strengthens model resilience.
Building fraud detection pipelines
Compliance and security teams can draw on these synthetic samples to assemble realistic training inputs without the regulatory risk of working with genuine documents. Because every passport field — including MRZ strings — is provided, both authentic-style records and suspicious anomalies can be modelled to support comprehensive fraud-detection workflows.
FAQ
Is this real-world or synthetic data?
All images are AI-generated and contain no biometric data or personal information tied to real individuals.
Can I request a custom dataset size?
Yes — the dataset is scalable, and additional samples can be generated based on your requirements within one week.
Can I request country-specific data?
Yes — support for 50+ countries is available. Please submit a request to get detailed coverage and samples.
Can I request a sample before purchasing?
Yes — free samples are available so you can evaluate image quality, metadata structure, and variation coverage before committing.
How is the dataset delivered?
After purchase, the dataset is delivered via secure AWS cloud infrastructure compliant with ISO 27001 and ISO 27701.
