datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic_us_passports_easy
Dataset Card for synthetic_us_passports
This is a FiftyOne dataset with 9750 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/synthetic_us_passports_easy")
# Launch the App
session = fo.launch_app(dataset)
Based on the… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/synthetic_us_passports_easy.synthetic-printed-japanese-passports
Japanese passport dataset
Dataset contains 5,000+ photos of synthetic Japanese passports, designed for training and validating Machine Learning models in PII extraction and document analysis. It features identity documents from a wide range of different countries and other countries, simulating the variety encountered in international travels.
By utilizing this synthetic dataset, researchers and businesses can advance their capabilities in biometric security, identity… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-japanese-passports.synthetic-printed-brazilian-passports
Brazilian passport dataset
The dataset comprises 5,000 high-resolution synthetic photos of Brazilian passports, designed to advance computer vision and identity verification systems. It provides a secure and ethical resource for training robust models for OCR (Optical Character Recognition), document analysis, and spoofing detection, all without exposing real personal data or sensitive personal information.
By utilizing this dataset, researchers and developers can enhance… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-brazilian-passports.synthetic-turkish-passports
Turkish passport dataset - 5, 000 images
Dataset comprises 5,000 meticulously organized files capturing Turkish passports under highly controlled variations, making it an invaluable resource for developing robust document recognition and verification systems. It is specifically designed for training and testing models in passport authentication, biometric data extraction, and identity verification.
By leveraging this dataset containing detailed information from Turkish… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-turkish-passports.synthetic-passports
Passport photos dataset
This dataset contains over 100,000 passport photos from 100+ countries, making it a valuable resource for researchers and developers working on computer vision tasks related to passport verification, biometric identification, and document analysis. This dataset allows researchers and developers to train and evaluate their models without the ethical and legal concerns associated with using real passport data.
By leveraging this dataset, developers can… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-passports.synthetic-printed-canadian-passports
Canadian passport dataset
Dataset includes 5,000 high-resolution, AI-generated passport images captured under varied angles, lighting, and backgrounds. Designed for OCR, computer vision, and identity verification research, this synthetic passport dataset provides diverse Canadian passport images for training secure document recognition and personal data extraction systems.
By utilizing this dataset, researchers and developers can train models to accurately read passport numbers… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-canadian-passports.synthetic-printed-australian-passports
Australian passport dataset
The dataset comprises 5,000 high-resolution synthetic photos of ** Australian passports**, designed to advance computer vision and identity verification systems. It provides a secure and ethical resource for training robust models for OCR (Optical Character Recognition), document analysis, and spoofing detection, all without exposing real personal data or sensitive personal information.
This dataset is an essential tool for organizations and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-australian-passports.synthetic_us_passports_easy
Synthetic US Passports (Hard)
This dataset is designed to evaluate VLMs transcription capabilities by using a well-known and straightforward document type: passports.
More specifically, it requires VLMs to be robust to:
tilted documents
high-resolution image with a small region of interest (since the passport only takes up a part of the image)
Note: there is a "sister" version of this dataset with some Augraphy augmentations (See:… See the full description on the dataset page: https://huggingface.co/datasets/arnaudstiegler/synthetic_us_passports_easy.generated-passports-segmentation
GENERATED USA Passports Segmentation
The dataset contains a collection of images representing GENERATED USA Passports. Each passport image is segmented into different zones, including the passport zone, photo, name, surname, date of birth, sex, nationality, passport number, and MRZ (Machine Readable Zone).
The dataset can be utilized for computer vision, object detection, data extraction and machine learning models.
Generated passports can assist in conducting research without… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/generated-passports-segmentation.passportsSource - https://www.dropbox.com/s/omintwb3k2h46kk/passport_dataset.zip
synthetic_us_passports_hard
Synthetic US Passports (Hard)
This dataset is designed to evaluate VLMs transcription capabilities by using a well-known and straightforward document type: passports.
More specifically, it requires VLMs to be robust to:
tilted documents
high-resolution image with a small region of interest (since the passport only takes up a part of the image)
HARD VERSION ONLY: noise injected using the Augraphy package, leading to a much more difficult transcription
Note: there is a "sister"… See the full description on the dataset page: https://huggingface.co/datasets/arnaudstiegler/synthetic_us_passports_hard.v2_synthetic_us_passports_easyv2_synthetic_us_passports_hardsynthetic-printed-nz-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset features 5,000 AI-generated New Zealand passport images captured under varied angles, lighting, and backgrounds. Designed for OCR, computer vision, and identity verification research, this NZ passport dataset supports training models in PII extraction, document recognition, and synthetic passport analysis with rich metadata annotations. - Get the data
Dataset characteristics:
Characteristic
Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-nz-passports.passportsSource - https://www.dropbox.com/s/omintwb3k2h46kk/passport_dataset.zip
synthetic-printed-japanese-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset contains 5,000 AI-generated, high-resolution passport images with diverse lighting, angles, and backgrounds. It supports document analysis, OCR, and biometric data research, offering realistic Japanese passport images for training and evaluating identity recognition and personal data extraction systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Printed synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-japanese-passports.synthetic-printed-uk-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset features 5,000 AI-generated UK passport images captured under varied angles, lighting, and backgrounds. Designed for OCR, computer vision, and identity verification research, this passport dataset supports training models in PII extraction, document recognition, and synthetic passport analysis with rich metadata annotations. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Printed… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-uk-passports.synthetic-printed-german-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset features 5,000 AI-generated German passport images captured under varied angles, lighting, and backgrounds. Designed for OCR, computer vision, and identity verification research, this passport dataset supports training models in PII extraction, document recognition, and synthetic passport analysis with rich metadata annotations. - Get the data
Dataset characteristics:
Characteristic
Data
Description… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-german-passports.synthetic-printed-canadian-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset provides 5,000 files with high-resolution synthetic passport images with diverse angles, lighting, and backgrounds, designed for training OCR, computer vision, and identity verification models without exposing real personal data or sensitive information. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Printed synthetic passport images for training ML models in PII extraction… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-canadian-passports.synthetic-turkish-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset containing 5,000 high-quality, AI-generated images. Labeled with detailed metadata - including passport ID, class, gender, and lighting - this dataset supports PII extraction, identity verification, and biometric recognition system training while maintaining strict data protection standards. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Printed synthetic passport images… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-turkish-passports.synthetic-printed-australian-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset provides 5,000 files with high-resolution synthetic passport images with diverse angles, lighting, and backgrounds, designed for training OCR, computer vision, and identity verification models without exposing real personal data or sensitive information. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Printed synthetic passport images for training ML models in PII extraction… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-australian-passports.synthetic-printed-brazilian-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset provides 5,000 files with high-resolution synthetic passport images with diverse angles, lighting, and backgrounds, designed for training OCR, computer vision, and identity verification models without exposing real personal data or sensitive information. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Printed synthetic passport images for training ML models in PII extraction… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-brazilian-passports.synthetic-printed-mexican-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset includes a diverse collection of AI-generated passport images replicating authentic Mexican passport layouts, fonts, and visual features. Designed for OCR, computer vision, and identity verification research, this passport dataset supports training models in PII extraction, document recognition, and synthetic passport analysis with rich metadata annotations. - Get the data
Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-mexican-passports.passport-stamps-problem
