CoolFace
Datasetpublic

ud-synthetic/brazilian-passports

Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction - Brazil The Synthetic Brazil Passports Dataset brings together over 1,000 AI-generated passport images designed for training OCR and computer vision systems on identity documents. Every record is fully synthetic, so the… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/brazilian-passports.

sourceHugging Facecc-by-nc-nd-4.0updated 2mo agoView on Hugging Face
1likes12downloads
Dataset Card
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.

Introduction - Brazil

The Synthetic Brazil Passports Dataset brings together over 1,000 AI-generated passport images designed for training OCR and computer vision systems on identity documents. Every record is fully synthetic, so the collection holds no real personal data or biometric details — making it a privacy-compliant resource for building identity verification, KYC, and fraud-detection pipelines. - [Get the data](https://unidata.pro/datasets/synthetic-passports/?utm_source=huggingface-synthetic&utm_medium=referral&utm_campaign=synthetic-brazil-passports)

The dataset is scalable on demand — extra images and metadata can be produced to your specifications, usually within one week.

Each image is captured on either a clean white background or a variety of realistic environments — desks, walls, and other everyday surfaces — so models adapt more reliably to production scenarios.

Coverage spans 50+ countries (including Mexico, Argentina, Colombia, Chile, Peru, and more). Submit a request via [the website](https://unidata.pro/datasets/synthetic-passports/?utm_source=huggingface-synthetic&utm_medium=referral&utm_campaign=synthetic-brazil-passports) to learn more.

Every image ships with rich structured metadata describing the document's personal fields — passport number, full name, signature, date of birth, sex, place of birth, issuing authority, nationality, document type, and the machine-readable zone (MRZ) — together with technical attributes such as resolution and category.


Dataset general info

CharacteristicData
DescriptionSynthetic passport images with detailed metadata for ML model training in PII extraction
Data typesImage + structured metadata
TasksOCR, Computer Vision
Total number of files1000+ (scalable on request)
LabelingPassport Number, Passport Number (split), Surname, Given Name, Signature, Date of Birth, Date of Issue, Date of Expiry, Sex, Place of Birth, Issuing Authority, Nationality, Nationality Code, Document Type, MRZ
GenderMale, Female
BackgroundsWhite and varied (desk, wall, and other surfaces)
Countries50+ available (Mexico, Argentina, Colombia, Chile, Peru, and more — on request)
Image formatJPG
Data generationAI-generated
Source of imagesAI-generated

The dataset comprises 1,000+ synthetic passport images, each tied to a complete identity-style record and structured annotations. Additional samples can be produced on request within one week, and the imagery covers both clean white scenes and varied background environments.

Metadata fields include:

FieldExample
Passport numberNN664353
Passport number (split)N N 6 6 4 3 5 3
Surname and given nameCASTRO GUSTAVO
SignatureGustavo Castro
Date of birth12 DEC 1965
Date of issue05 NOV 2023
Date of expiry05 NOV 2033
SexM
Place of birthSAO PAULO/SP
Issuing authoritySR/PF/MG
NationalityBRAZIL
Nationality codeBRA
Document typeP
Machine-readable zone (MRZ)P<BRACASTRO<<GUSTAVO<<<<<<<<<<<<<<<<<<<<<<<<NN664353<7BRA6512127M3311053<<<<<<<<<<<<<<04

Use cases - Brazil

Training document verification systems

Banks, fintech operators, and border-control teams can apply this dataset to train models that verify identity documents across a wide range of real-world conditions. The structured metadata supports accurate field extraction and validation, while the variety of backgrounds strengthens robustness in production deployments.

Building fraud detection pipelines

Compliance and security teams can use these synthetic samples to compose high-quality training corpora without relying on actual travel documents. Complete passport field coverage and full MRZ strings make it possible to model and test both standard and anomalous record patterns with confidence.


FAQ

Is this real-world or synthetic data?

All images are AI-generated and contain no biometric data or personal information tied to real individuals.

Can I request a custom dataset size?

Yes — the dataset is scalable, and additional samples can be generated based on your requirements within one week.

Can I request country-specific data?

Yes — support for 50+ countries is available. Please submit a request to get detailed coverage and samples.

Can I request a sample before purchasing?

Yes — free samples are available so you can evaluate image quality, metadata structure, and variation coverage before committing.

How is the dataset delivered?

After purchase, the dataset is delivered via secure AWS cloud infrastructure compliant with ISO 27001 and ISO 27701.

🌐 UniData - your trusted data partner. Unique, accurate, thoroughly collected and annotated data designed to fuel your AI/ML success.