CoolFace
Datasetpublic

ud-synthetic/japanese-passports

Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction The Synthetic Japan Passports Dataset brings together more than 1,000 AI-generated passport images, purpose-built for training OCR and computer vision models on identity documents. Since the data is entirely synthetic —… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/japanese-passports.

sourceHugging Facecc-by-nc-nd-4.0updated 2mo agoView on Hugging Face
1likes78downloads
Dataset Card
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.

Introduction

The Synthetic Japan Passports Dataset brings together more than 1,000 AI-generated passport images, purpose-built for training OCR and computer vision models on identity documents. Since the data is entirely synthetic — with no genuine personal records or biometric details involved — it provides a privacy-safe option for teams developing identity verification and fraud detection solutions. - [Get the data](https://unidata.pro/datasets/synthetic-passports/?utm_source=huggingface-synthetic&utm_medium=referral&utm_campaign=synthetic-japan-passports)

The dataset can be expanded to fit your needs — additional images and metadata can be generated to your specifications, with new samples typically ready within one week.

Imagery covers both clean white backgrounds and a variety of realistic everyday scenes, including desks, walls, and other common surfaces, giving models the visual diversity needed to perform well in production environments.

The dataset spans 50+ countries (including China, South Korea, Singapore, Indonesia, Vietnam, and more). To learn more, please submit a request via [the website](https://unidata.pro/datasets/synthetic-passports/?utm_source=huggingface-synthetic&utm_medium=referral&utm_campaign=synthetic-japan-passports) to learn more.

Each image comes with comprehensive structured metadata that captures personal document fields — passport number, full name, signature, date of birth, sex, place of birth, issuing authority, nationality, document type, and machine-readable zone (MRZ) — together with technical attributes such as resolution and category.


Dataset general info

CharacteristicData
DescriptionSynthetic passport images with detailed metadata for ML model training in PII extraction
Data typesImage + structured metadata
TasksOCR, Computer Vision
Total number of files1000+ (scalable on request)
LabelingPassport Number, Passport Number (split), Surname, Given Name, Signature, Date of Birth, Date of Issue, Date of Expiry, Sex, Place of Birth, Issuing Authority, Nationality, Nationality Code, Document Type, MRZ
GenderMale, Female
BackgroundsWhite and varied (desk, wall, and other surfaces)
Countries50+ available (China, South Korea, Singapore, Indonesia, Vietnam, and more — on request)
Image formatJPG
Data generationAI-generated
Source of imagesAI-generated

The dataset consists of 1000+ synthetic passport images, each associated with complete identity-like records and structured annotations. Additional data can be generated upon request within one week. The images cover both clean white and diverse background scenarios.

Metadata fields include:

FieldExample
Passport numberQI7859558
Passport number (split)Q I 7 8 5 9 5 5 8
Surname and given nameYAMAGUCHI AKIRA
SignatureAkira Yamaguchi
Date of birth26 SEP 1968
Date of issue06 OCT 2020
Date of expiry06 OCT 2030
SexM
Place of birthTOKYO
Issuing authorityFUKUOKA
NationalityJAPAN
Nationality codeJPN
Document typeP
Machine-readable zone (MRZ)P<JPNYAMAGUCHI<<AKIRA<<<<<<<<<<<<<<<<<<<<<<<

Use cases

Training document verification systems

Banks, fintech operators, and immigration authorities can apply this dataset to develop models that authenticate identity documents under varied real-world conditions. The structured metadata enables accurate field-level extraction and validation, while the assortment of background settings strengthens model resilience.

Building fraud detection pipelines

Compliance and security teams can draw on these synthetic samples to assemble realistic training inputs without the regulatory risk of working with genuine documents. Because every passport field — including MRZ strings — is provided, both authentic-style records and suspicious anomalies can be modelled to support comprehensive fraud-detection workflows.


FAQ

Is this real-world or synthetic data?

All images are AI-generated and contain no biometric data or personal information tied to real individuals.

Can I request a custom dataset size?

Yes — the dataset is scalable, and additional samples can be generated based on your requirements within one week.

Can I request country-specific data?

Yes — support for 50+ countries is available. Please submit a request to get detailed coverage and samples.

Can I request a sample before purchasing?

Yes — free samples are available so you can evaluate image quality, metadata structure, and variation coverage before committing.

How is the dataset delivered?

After purchase, the dataset is delivered via secure AWS cloud infrastructure compliant with ISO 27001 and ISO 27701.

🌐 UniData - your trusted data partner. Unique, accurate, thoroughly collected and annotated data designed to fuel your AI/ML success.