datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-printed-ocr-3.5m
Persian Printed OCR 3.5M
A unified corpus of 3,517,974 Persian printed OCR image/text pairs, selected from five public
datasets using GlotLID v3. Only the accept bucket is included; 232,317 ambiguous and 190,733
rejected rows are excluded. The viewer exposes exactly image and label.
Sources
AliShafiee2003/persian-ocr-garshasp-70c — pinned revision 36bfdcdeac20c02231f4ee08472f80db2fc467bb (CC-BY-4.0)
hezarai/parsynth-ocr-200k — pinned revision… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-printed-ocr-3.5m.early_printed_books_font_detection
Early Printed Books Font Detection
Photographs of 35,623 pages from books printed between the mid-15th and the end of the 18th century, each labelled by experts with the font group or groups used on the page. This is a mirror of Dataset of Pages from Early Printed Books with Multiple Font Groups by Mathias Seuret, Saskia Limbach, Nikolaus Weichselbaumer, Andreas Maier and Vincent Christlein, deposited on Zenodo in August 2019 and described in their HIP'19 paper.
The page images… See the full description on the dataset page: https://huggingface.co/datasets/biglam/early_printed_books_font_detection.synthetic-printed-japanese-passports
Japanese passport dataset
Dataset contains 5,000+ photos of synthetic Japanese passports, designed for training and validating Machine Learning models in PII extraction and document analysis. It features identity documents from a wide range of different countries and other countries, simulating the variety encountered in international travels.
By utilizing this synthetic dataset, researchers and businesses can advance their capabilities in biometric security, identity… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-japanese-passports.synthetic-printed-brazilian-passports
Brazilian passport dataset
The dataset comprises 5,000 high-resolution synthetic photos of Brazilian passports, designed to advance computer vision and identity verification systems. It provides a secure and ethical resource for training robust models for OCR (Optical Character Recognition), document analysis, and spoofing detection, all without exposing real personal data or sensitive personal information.
By utilizing this dataset, researchers and developers can enhance… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-brazilian-passports.synthetic-printed-canadian-passports
Canadian passport dataset
Dataset includes 5,000 high-resolution, AI-generated passport images captured under varied angles, lighting, and backgrounds. Designed for OCR, computer vision, and identity verification research, this synthetic passport dataset provides diverse Canadian passport images for training secure document recognition and personal data extraction systems.
By utilizing this dataset, researchers and developers can train models to accurately read passport numbers… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-canadian-passports.early_printed_books_font_detection_loaded
Dataset Card for "early_printed_books_font_detection_loaded"
More Information needed
OCR-English-Printed-12A synthetic dataset for text recognition tasks, contains 1.000.000 images
ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz
62k-images-khmer-printed-dataset
62k Khmer-English Printed Dataset
This repository contains a dataset of Khmer and English printed text images for training, validation, and testing. The dataset is stored in parquet format and managed using Git Large File Storage (LFS).
Installation
Prerequisites
Before cloning this repository, make sure you have Git LFS installed:
Install Git LFS
Linux/macOS:curl -s https://packagecloud.io/install/repositories/github/git-lfs/script.deb.sh | sudo… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/62k-images-khmer-printed-dataset.synthetic-printed-australian-passports
Australian passport dataset
The dataset comprises 5,000 high-resolution synthetic photos of ** Australian passports**, designed to advance computer vision and identity verification systems. It provides a secure and ethical resource for training robust models for OCR (Optical Character Recognition), document analysis, and spoofing detection, all without exposing real personal data or sensitive personal information.
This dataset is an essential tool for organizations and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-australian-passports.3d-printed-or-not
3d-printed-or-not: An Image Dataset of 3D-printed Prototypes
This dataset is a collection of images that are particularly relevant to engineering and design, consisting of two categories: 3D-printed prototypes, and non-3D-printed prototypes This data was collected through a hybrid approach that entailed both web scraping and direct collection from engineering labs and workspaces at Penn State University. The initial data was then augmented using several data augmentation techniques… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/3d-printed-or-not.rus_xviii_printed_geography
OCR Dataset of 18th-Century Russian Printed Texts
Dataset Overview
This dataset contains line-level image–text pairs extracted from historical Russian printed materials of the 18th century. It is designed for training and evaluating Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) systems on pre-reform Russian orthography.
The dataset is intended for research in computational linguistics, digital humanities, and historical text… See the full description on the dataset page: https://huggingface.co/datasets/dsmchr/rus_xviii_printed_geography.OCR-Cyrillic-Printed-10A synthetic dataset for text recognition tasks, contains 1.000.000 images
АБВГДЕЁЖЗИЙКЛМНОПРСТУФХЦЧШЩЪЫЬЭЮЯабвгдеёжзийклмнопрстуфхцчшщъыьэюя
synthetic-printed-nz-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset features 5,000 AI-generated New Zealand passport images captured under varied angles, lighting, and backgrounds. Designed for OCR, computer vision, and identity verification research, this NZ passport dataset supports training models in PII extraction, document recognition, and synthetic passport analysis with rich metadata annotations. - Get the data
Dataset characteristics:
Characteristic
Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-nz-passports.OCR-Cyrillic-Printed-8A synthetic dataset for text recognition tasks, contains 1.000.000 images
АБВГДЕЁЖЗИЙКЛМНОПРСТУФХЦЧШЩЪЫЬЭЮЯабвгдеёжзийклмнопрстуфхцчшщъыьэюя
printed-circuit-board
Dataset Card for printed-circuit-board
** The original COCO dataset is stored at dataset.tar.gz**
Dataset Summary
printed-circuit-board
Supported Tasks and Leaderboards
object-detection: The dataset can be used to train a model for Object Detection.
Languages
English
Dataset Structure
Data Instances
A data point comprises an image and its object annotations.
{
'image_id': 15,
'image': <PIL.JpegImagePlugin.JpegImageFile… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/printed-circuit-board.OCR-Cyrillic-Printed-9A synthetic dataset for text recognition tasks, contains 300,000 images
АБВГДЕЁЖЗИЙКЛМНОПРСТУФХЦЧШЩЪЫЬЭЮЯабвгдеёжзийклмнопрстуфхцчшщъыьэюя,./:;''"[]{}-_=+!?*()~<>^\«»#
synthetic-printed-japanese-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset contains 5,000 AI-generated, high-resolution passport images with diverse lighting, angles, and backgrounds. It supports document analysis, OCR, and biometric data research, offering realistic Japanese passport images for training and evaluating identity recognition and personal data extraction systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Printed synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-japanese-passports.bangla-ocr-validation_data_printed
Bangla OCR Validation Dataset (Printed + Scanned)
📌 Description
This dataset is a Bangla OCR validation dataset containing a mix of printed document images and their corresponding text annotations. It is designed to evaluate OCR and vision-language models on both clean digital text and scanned document images.
📊 Dataset Composition
1507 line-level images with text annotations
50 full-page document images with text
Data includes:
Printed/typed Bangla text… See the full description on the dataset page: https://huggingface.co/datasets/arobin79/bangla-ocr-validation_data_printed.kazakh-printed-dataset
Kazakh Printed Dataset for OCR task
Data Lineage
This dataset was synthetically generated using issai/kazparc as the base.
Since Kazakh OCR data is scarce, I developed a pipeline to transform digital Kazakh text into a printed-style dataset.
Generation Process
Source:
Text samples were extracted from issai/kazparc.
Augmentation & Stylization:
Random Background Color: Simulates different lighting conditions by alternating between… See the full description on the dataset page: https://huggingface.co/datasets/thekamilya/kazakh-printed-dataset.synthetic-printed-uk-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset features 5,000 AI-generated UK passport images captured under varied angles, lighting, and backgrounds. Designed for OCR, computer vision, and identity verification research, this passport dataset supports training models in PII extraction, document recognition, and synthetic passport analysis with rich metadata annotations. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Printed… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-uk-passports.synthetic-printed-german-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset features 5,000 AI-generated German passport images captured under varied angles, lighting, and backgrounds. Designed for OCR, computer vision, and identity verification research, this passport dataset supports training models in PII extraction, document recognition, and synthetic passport analysis with rich metadata annotations. - Get the data
Dataset characteristics:
Characteristic
Data
Description… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-german-passports.synthetic-printed-canadian-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset provides 5,000 files with high-resolution synthetic passport images with diverse angles, lighting, and backgrounds, designed for training OCR, computer vision, and identity verification models without exposing real personal data or sensitive information. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Printed synthetic passport images for training ML models in PII extraction… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-canadian-passports.OCR-Cyrillic-Printed-6A synthetic dataset for text recognition tasks, contains 1.000.000 images
АБВГДЕЁЖЗИЙКЛМНОПРСТУФХЦЧШЩЪЫЬЭЮЯабвгдеёжзийклмнопрстуфхцчшщъыьэюя
synthetic-printed-australian-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset provides 5,000 files with high-resolution synthetic passport images with diverse angles, lighting, and backgrounds, designed for training OCR, computer vision, and identity verification models without exposing real personal data or sensitive information. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Printed synthetic passport images for training ML models in PII extraction… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-australian-passports.OCR-Cyrillic-Printed-1A synthetic dataset for text recognition tasks, contains 500,000 images
OCR-English-Printed-14A synthetic dataset for text recognition tasks, contains 1.000.000 images
ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz
synthetic-printed-brazilian-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset provides 5,000 files with high-resolution synthetic passport images with diverse angles, lighting, and backgrounds, designed for training OCR, computer vision, and identity verification models without exposing real personal data or sensitive information. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Printed synthetic passport images for training ML models in PII extraction… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-brazilian-passports.OCR-Numbers-Printed-0A synthetic dataset for text recognition tasks, contains 300,000 images, numbers only.
synthetic-printed-mexican-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset includes a diverse collection of AI-generated passport images replicating authentic Mexican passport layouts, fonts, and visual features. Designed for OCR, computer vision, and identity verification research, this passport dataset supports training models in PII extraction, document recognition, and synthetic passport analysis with rich metadata annotations. - Get the data
Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-mexican-passports.OCR-Numbers-Printed-5A synthetic dataset for text recognition tasks, contains 100,000 images, numbers only.
