CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Reza2kn /persian-printed-ocr-3.5m Persian Printed OCR 3.5M A unified corpus of 3,517,974 Persian printed OCR image/text pairs, selected from five public datasets using GlotLID v3. Only the accept bucket is included; 232,317 ambiguous and 190,733 rejected rows are excluded. The viewer exposes exactly image and label. Sources AliShafiee2003/persian-ocr-garshasp-70c — pinned revision 36bfdcdeac20c02231f4ee08472f80db2fc467bb (CC-BY-4.0) hezarai/parsynth-ocr-200k — pinned revision… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-printed-ocr-3.5m.imageimage-to-text1M<n<10M2 likes1.5k downloads2mo agoHugging Face02biglam /early_printed_books_font_detection Early Printed Books Font Detection Photographs of 35,623 pages from books printed between the mid-15th and the end of the 18th century, each labelled by experts with the font group or groups used on the page. This is a mirror of Dataset of Pages from Early Printed Books with Multiple Font Groups by Mathias Seuret, Saskia Limbach, Nikolaus Weichselbaumer, Andreas Maier and Vincent Christlein, deposited on Zenodo in August 2019 and described in their HIP'19 paper. The page images… See the full description on the dataset page: https://huggingface.co/datasets/biglam/early_printed_books_font_detection.imageimage-classification10K<n<100K2 likes425 downloads2mo agoHugging Face03UniDataPro /synthetic-printed-japanese-passports Japanese passport dataset Dataset contains 5,000+ photos of synthetic Japanese passports, designed for training and validating Machine Learning models in PII extraction and document analysis. It features identity documents from a wide range of different countries and other countries, simulating the variety encountered in international travels. By utilizing this synthetic dataset, researchers and businesses can advance their capabilities in biometric security, identity… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-japanese-passports.imageimage-to-textn<1K0 likes397 downloads1mo agoHugging Face04UniDataPro /synthetic-printed-brazilian-passports Brazilian passport dataset The dataset comprises 5,000 high-resolution synthetic photos of Brazilian passports, designed to advance computer vision and identity verification systems. It provides a secure and ethical resource for training robust models for OCR (Optical Character Recognition), document analysis, and spoofing detection, all without exposing real personal data or sensitive personal information. By utilizing this dataset, researchers and developers can enhance… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-brazilian-passports.imageimage-to-textn<1K1 likes360 downloads1mo agoHugging Face05LibreYOLO /printed-circuit-board Printed Circuit Board This dataset is part of the Roboflow 100 benchmark, a diverse collection of 100 object detection datasets spanning 7 imagery domains. Dataset Statistics Split Images Train 548 Validation 80 Test 44 Total 672 Classes (34) Battery Button Buzzer Capacitor Jumper Capacitor Clock Connector Diode Display EM Electrolytic Capacitor Ferrite Bead Fuse Heatsink IC Inductor Jumper Led PS Pads Pins Potentiometer Resistor… See the full description on the dataset page: https://huggingface.co/datasets/LibreYOLO/printed-circuit-board.object-detection1K<n<10K0 likes329 downloads8mo agoHugging Face06UniDataPro /synthetic-printed-canadian-passports Canadian passport dataset Dataset includes 5,000 high-resolution, AI-generated passport images captured under varied angles, lighting, and backgrounds. Designed for OCR, computer vision, and identity verification research, this synthetic passport dataset provides diverse Canadian passport images for training secure document recognition and personal data extraction systems. By utilizing this dataset, researchers and developers can train models to accurately read passport numbers… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-canadian-passports.imageimage-to-textn<1K0 likes226 downloads1mo agoHugging Face07biglam /early_printed_books_font_detection_loaded Dataset Card for "early_printed_books_font_detection_loaded" More Information needed image1K<n<10K0 likes218 downloads4y agoHugging Face08DonkeySmall /OCR-English-Printed-12A synthetic dataset for text recognition tasks, contains 1.000.000 images ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz imageimage-to-text1M<n<10M3 likes216 downloads2y agoHugging Face09SoyVitou /62k-images-khmer-printed-dataset 62k Khmer-English Printed Dataset This repository contains a dataset of Khmer and English printed text images for training, validation, and testing. The dataset is stored in parquet format and managed using Git Large File Storage (LFS). Installation Prerequisites Before cloning this repository, make sure you have Git LFS installed: Install Git LFS Linux/macOS:curl -s https://packagecloud.io/install/repositories/github/git-lfs/script.deb.sh | sudo… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/62k-images-khmer-printed-dataset.imagetext-generation10K<n<100K2 likes199 downloads2y agoHugging Face10UniDataPro /synthetic-printed-australian-passports Australian passport dataset The dataset comprises 5,000 high-resolution synthetic photos of ** Australian passports**, designed to advance computer vision and identity verification systems. It provides a secure and ethical resource for training robust models for OCR (Optical Character Recognition), document analysis, and spoofing detection, all without exposing real personal data or sensitive personal information. This dataset is an essential tool for organizations and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-australian-passports.imageimage-to-textn<1K0 likes174 downloads1mo agoHugging Face11UniDataPro /synthetic-printed-usa-passports-dataset Passport Dataset - 9 600 Images The dataset comprises 9,600 high-quality synthetically generated passport images, providing a robust resource for training and verifying document analysis systems. Every passport is presented across 3 angles (0°, 25°, 45°), 4 lighting conditions (Natural-daylight, Office-LED, Warm-indoor, Dim-light), 4 backgrounds (Neutral wall, Textured desk, Outdoor pavement, Docs-on-docs), and 2 distances (Close, Medium), creating a rich and challenging dataset… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-usa-passports-dataset.image-to-text1K<n<10K1 likes144 downloads1mo agoHugging Face12cmudrc /3d-printed-or-not 3d-printed-or-not: An Image Dataset of 3D-printed Prototypes This dataset is a collection of images that are particularly relevant to engineering and design, consisting of two categories: 3D-printed prototypes, and non-3D-printed prototypes This data was collected through a hybrid approach that entailed both web scraping and direct collection from engineering labs and workspaces at Penn State University. The initial data was then augmented using several data augmentation techniques… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/3d-printed-or-not.imageimage-classification10K<n<100K1 likes104 downloads4y agoHugging Face13dsmchr /rus_xviii_printed_geography OCR Dataset of 18th-Century Russian Printed Texts Dataset Overview This dataset contains line-level image–text pairs extracted from historical Russian printed materials of the 18th century. It is designed for training and evaluating Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) systems on pre-reform Russian orthography. The dataset is intended for research in computational linguistics, digital humanities, and historical text… See the full description on the dataset page: https://huggingface.co/datasets/dsmchr/rus_xviii_printed_geography.image1K<n<10K0 likes103 downloads3mo agoHugging Face14UniqueData /printed-2d-masks-with-holes-for-eyes-attacksThe dataset consists of selfies of people and videos of them wearing a printed 2d mask with their face. The dataset solves tasks in the field of anti-spoofing and it is useful for buisness and safety systems. The dataset includes: **attacks** - videos of people wearing printed portraits of themselves with cut-out eyes.image-classification10K<n<100K2 likes69 downloads1y agoHugging Face15DonkeySmall /OCR-Cyrillic-Printed-10A synthetic dataset for text recognition tasks, contains 1.000.000 images АБВГДЕЁЖЗИЙКЛМНОПРСТУФХЦЧШЩЪЫЬЭЮЯабвгдеёжзийклмнопрстуфхцчшщъыьэюя imageimage-to-text1M<n<10M0 likes69 downloads2y agoHugging Face16PANDITSAGAR /printed-circuit-board Printed Circuit Board This dataset is part of the Roboflow 100 benchmark, a diverse collection of 100 object detection datasets spanning 7 imagery domains. Dataset Statistics Split Images Train 548 Validation 80 Test 44 Total 672 Classes (34) Battery Button Buzzer Capacitor Jumper Capacitor Clock Connector Diode Display EM Electrolytic Capacitor Ferrite Bead Fuse Heatsink IC Inductor Jumper Led PS Pads Pins Potentiometer… See the full description on the dataset page: https://huggingface.co/datasets/PANDITSAGAR/printed-circuit-board.object-detection1K<n<10K0 likes69 downloads21d agoHugging Face17UniqueData /printed_photos_attacksThe dataset consists of 40,000 videos and selfies with unique people. 15,000 attack replays from 4,000 unique devices. 10,000 attacks with A4 printouts and 10,000 attacks with cut-out printouts.image-to-image10K<n<100K1 likes64 downloads1y agoHugging Face18ud-biometrics /synthetic-printed-nz-passports Synthetic Passports Dataset - 5 000 passport photos Dataset features 5,000 AI-generated New Zealand passport images captured under varied angles, lighting, and backgrounds. Designed for OCR, computer vision, and identity verification research, this NZ passport dataset supports training models in PII extraction, document recognition, and synthetic passport analysis with rich metadata annotations. - Get the data Dataset characteristics: Characteristic Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-nz-passports.imageimage-to-textn<1K0 likes57 downloads11mo agoHugging Face19UniqueData /attacks-with-2d-printed-masks-of-indian-people Attacks with 2D Printed Masks of Indian People - Biometric Attack Dataset The dataset consists of videos of individuals wearing printed 2D masks of different kinds and directly looking at the camera. Videos are filmed in different lightning conditions and in different places (indoors, outdoors). Each video in the dataset has an approximate duration of 3-4 seconds. The similar dataset that includes all ethnicities - Printed 2D Masks Attacks Dataset Types of… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/attacks-with-2d-printed-masks-of-indian-people.videovideo-classificationn<1K1 likes51 downloads1y agoHugging Face20DonkeySmall /OCR-Cyrillic-Printed-8A synthetic dataset for text recognition tasks, contains 1.000.000 images АБВГДЕЁЖЗИЙКЛМНОПРСТУФХЦЧШЩЪЫЬЭЮЯабвгдеёжзийклмнопрстуфхцчшщъыьэюя imageimage-to-text100K<n<1M0 likes50 downloads2y agoHugging Face21UniqueData /2d-printed_masks_attacksThe dataset consists of 40,000 videos and selfies with unique people. 15,000 attack replays from 4,000 unique devices. 10,000 attacks with A4 printouts and 10,000 attacks with cut-out printouts.video-classification1K<n<10K1 likes46 downloads1y agoHugging Face22medyas /arabic-ocr-printed-500k Arabic Printed OCR Lines — Synthetic, 500k A general-purpose printed Arabic text-line recognition corpus: 500,000 train + 2,000 val line images with labels, built to fine-tune line-recognition models (PaddleOCR PP-OCR rec CTC/MultiHead, TrOCR, etc.). Real line-crop printed-Arabic data does not exist at this scale on the Hub, so this corpus is rendered synthetically with diverse fonts + real Arabic text and a documented label/decoding contract. Why this exists… See the full description on the dataset page: https://huggingface.co/datasets/medyas/arabic-ocr-printed-500k.tabularimage-to-textn<1K0 likes43 downloads3mo agoHugging Face23Francesco /printed-circuit-board Dataset Card for printed-circuit-board ** The original COCO dataset is stored at dataset.tar.gz** Dataset Summary printed-circuit-board Supported Tasks and Leaderboards object-detection: The dataset can be used to train a model for Object Detection. Languages English Dataset Structure Data Instances A data point comprises an image and its object annotations. { 'image_id': 15, 'image': <PIL.JpegImagePlugin.JpegImageFile… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/printed-circuit-board.imageobject-detectionn<1K0 likes39 downloads3y agoHugging Face24UniDataPro /2d-printed-mask-dataset 2D Mask Attack Dataset - 26 436 videos The dataset comprises 26,436 videos of real faces, 2D print attacks (printed photos), and replay attacks (faces displayed on screens), captured under varied conditions. Designed for attack detection research, it supports the development of robust face antispoofing and spoofing detection methods, critical for facial recognition security. Ideal for training models and refining anti-spoofing methods, the dataset enhances detection accuracy in… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/2d-printed-mask-dataset.videovideo-classificationn<1K1 likes39 downloads1mo agoHugging Face25UniDataPro /printed-2d-masks-attacks 2D Masks Attack for facial recogniton system The dataset consists of 4,800+ videos of people wearing of holding 2D printed masks filmed using 5 devices. It is designed for liveness detection algorithms, specifically aimed at enhancing anti-spoofing capabilities in biometric security systems. By leveraging this dataset, researchers can create more sophisticated recognition system, crucial for achieving iBeta Level 1 & 2 certification – a key standard for secure and reliable… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/printed-2d-masks-attacks.videovideo-classificationn<1K1 likes38 downloads1mo agoHugging Face26Akshit03 /AkshitMajorProjectMIR1_strawberry_printed_session2_20260721_125019This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Akshit03/AkshitMajorProjectMIR1_strawberry_printed_session2_20260721_125019.tabularrobotics1K<n<10K0 likes32 downloads2mo agoHugging Face27Akshit03 /AkshitMajorProjectMIR1_strawberry_printed_session3_20260721_132013This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Akshit03/AkshitMajorProjectMIR1_strawberry_printed_session3_20260721_132013.tabularrobotics1K<n<10K0 likes32 downloads2mo agoHugging Face28DonkeySmall /OCR-Cyrillic-Printed-9A synthetic dataset for text recognition tasks, contains 300,000 images АБВГДЕЁЖЗИЙКЛМНОПРСТУФХЦЧШЩЪЫЬЭЮЯабвгдеёжзийклмнопрстуфхцчшщъыьэюя,./:;''"[]{}-_=+!?*()~<>^\«»# image100K<n<1M1 likes31 downloads2y agoHugging Face29ud-biometrics /synthetic-printed-japanese-passports Synthetic Passports Dataset - 5 000 passport photos Dataset contains 5,000 AI-generated, high-resolution passport images with diverse lighting, angles, and backgrounds. It supports document analysis, OCR, and biometric data research, offering realistic Japanese passport images for training and evaluating identity recognition and personal data extraction systems. - Get the data Dataset characteristics: Characteristic Data Description Printed synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-japanese-passports.imageimage-to-textn<1K0 likes31 downloads11mo agoHugging Face30arobin79 /bangla-ocr-validation_data_printed Bangla OCR Validation Dataset (Printed + Scanned) 📌 Description This dataset is a Bangla OCR validation dataset containing a mix of printed document images and their corresponding text annotations. It is designed to evaluate OCR and vision-language models on both clean digital text and scanned document images. 📊 Dataset Composition 1507 line-level images with text annotations 50 full-page document images with text Data includes: Printed/typed Bangla text… See the full description on the dataset page: https://huggingface.co/datasets/arobin79/bangla-ocr-validation_data_printed.imageimage-to-text1K<n<10K1 likes30 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.