datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-ocr-printed-500k
Arabic Printed OCR Lines — Synthetic, 500k
A general-purpose printed Arabic text-line recognition corpus: 500,000 train + 2,000 val
line images with labels, built to fine-tune line-recognition models (PaddleOCR PP-OCR rec
CTC/MultiHead, TrOCR, etc.). Real line-crop printed-Arabic data does not exist at this scale
on the Hub, so this corpus is rendered synthetically with diverse fonts + real Arabic text and
a documented label/decoding contract.
Why this exists… See the full description on the dataset page: https://huggingface.co/datasets/medyas/arabic-ocr-printed-500k.printed-usa-passports
Introduction
The Synthetic Printed USA Passports Dataset contains 9,600 AI-generated passport images designed for training OCR and computer vision models on identity documents. The dataset includes varied angles, lighting conditions, backgrounds, and distances, with structured metadata covering gender, age group, resolution, and more. All images are synthetically generated — no real personal data or biometric records are involved — making it a privacy-compliant solution for… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/printed-usa-passports.printed-german-passports
Introduction
The Synthetic Printed German Passports Dataset contains 5,000 AI-generated passport images built for training OCR and computer vision models on printed identification documents. Each image is captured across 3 angles, 4 lighting conditions, 4 backgrounds, and 2 distances, with structured metadata covering passport ID, gender, age group, and more. Since all images are synthetically generated, the dataset contains no real personal data or biometric records — making it… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/printed-german-passports.
