printed
Datasets
All datasets matching “printed”persian-printed-ocr-3.5m
Persian Printed OCR 3.5M
A unified corpus of 3,517,974 Persian printed OCR image/text pairs, selected from five public
datasets using GlotLID v3. Only the accept bucket is included; 232,317 ambiguous and 190,733
rejected rows are excluded. The viewer exposes exactly image and label.
Sources
AliShafiee2003/persian-ocr-garshasp-70c — pinned revision 36bfdcdeac20c02231f4ee08472f80db2fc467bb (CC-BY-4.0)
hezarai/parsynth-ocr-200k — pinned revision… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-printed-ocr-3.5m.early_printed_books_font_detection
Early Printed Books Font Detection
Photographs of 35,623 pages from books printed between the mid-15th and the end of the 18th century, each labelled by experts with the font group or groups used on the page. This is a mirror of Dataset of Pages from Early Printed Books with Multiple Font Groups by Mathias Seuret, Saskia Limbach, Nikolaus Weichselbaumer, Andreas Maier and Vincent Christlein, deposited on Zenodo in August 2019 and described in their HIP'19 paper.
The page images… See the full description on the dataset page: https://huggingface.co/datasets/biglam/early_printed_books_font_detection.synthetic-printed-japanese-passports
Japanese passport dataset
Dataset contains 5,000+ photos of synthetic Japanese passports, designed for training and validating Machine Learning models in PII extraction and document analysis. It features identity documents from a wide range of different countries and other countries, simulating the variety encountered in international travels.
By utilizing this synthetic dataset, researchers and businesses can advance their capabilities in biometric security, identity… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-japanese-passports.synthetic-printed-brazilian-passports
Brazilian passport dataset
The dataset comprises 5,000 high-resolution synthetic photos of Brazilian passports, designed to advance computer vision and identity verification systems. It provides a secure and ethical resource for training robust models for OCR (Optical Character Recognition), document analysis, and spoofing detection, all without exposing real personal data or sensitive personal information.
By utilizing this dataset, researchers and developers can enhance… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-brazilian-passports.printed-circuit-board
Printed Circuit Board
This dataset is part of the Roboflow 100 benchmark, a diverse collection of 100 object detection datasets spanning 7 imagery domains.
Dataset Statistics
Split
Images
Train
548
Validation
80
Test
44
Total
672
Classes (34)
Battery
Button
Buzzer
Capacitor Jumper
Capacitor
Clock
Connector
Diode
Display
EM
Electrolytic Capacitor
Ferrite Bead
Fuse
Heatsink
IC
Inductor
Jumper
Led
PS
Pads
Pins
Potentiometer
Resistor… See the full description on the dataset page: https://huggingface.co/datasets/LibreYOLO/printed-circuit-board.synthetic-printed-canadian-passports
Canadian passport dataset
Dataset includes 5,000 high-resolution, AI-generated passport images captured under varied angles, lighting, and backgrounds. Designed for OCR, computer vision, and identity verification research, this synthetic passport dataset provides diverse Canadian passport images for training secure document recognition and personal data extraction systems.
By utilizing this dataset, researchers and developers can train models to accurately read passport numbers… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-canadian-passports.
