CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Reza2kn /persian-printed-ocr-3.5m Persian Printed OCR 3.5M A unified corpus of 3,517,974 Persian printed OCR image/text pairs, selected from five public datasets using GlotLID v3. Only the accept bucket is included; 232,317 ambiguous and 190,733 rejected rows are excluded. The viewer exposes exactly image and label. Sources AliShafiee2003/persian-ocr-garshasp-70c — pinned revision 36bfdcdeac20c02231f4ee08472f80db2fc467bb (CC-BY-4.0) hezarai/parsynth-ocr-200k — pinned revision… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-printed-ocr-3.5m.imageimage-to-text1M<n<10M2 likes1.5k downloads2mo agoHugging Face02biglam /early_printed_books_font_detection Early Printed Books Font Detection Photographs of 35,623 pages from books printed between the mid-15th and the end of the 18th century, each labelled by experts with the font group or groups used on the page. This is a mirror of Dataset of Pages from Early Printed Books with Multiple Font Groups by Mathias Seuret, Saskia Limbach, Nikolaus Weichselbaumer, Andreas Maier and Vincent Christlein, deposited on Zenodo in August 2019 and described in their HIP'19 paper. The page images… See the full description on the dataset page: https://huggingface.co/datasets/biglam/early_printed_books_font_detection.imageimage-classification10K<n<100K2 likes358 downloads2mo agoHugging Face03biglam /early_printed_books_font_detection_loaded Dataset Card for "early_printed_books_font_detection_loaded" More Information needed image1K<n<10K0 likes218 downloads4y agoHugging Face04SoyVitou /62k-images-khmer-printed-dataset 62k Khmer-English Printed Dataset This repository contains a dataset of Khmer and English printed text images for training, validation, and testing. The dataset is stored in parquet format and managed using Git Large File Storage (LFS). Installation Prerequisites Before cloning this repository, make sure you have Git LFS installed: Install Git LFS Linux/macOS:curl -s https://packagecloud.io/install/repositories/github/git-lfs/script.deb.sh | sudo… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/62k-images-khmer-printed-dataset.imagetext-generation10K<n<100K2 likes182 downloads2y agoHugging Face05cmudrc /3d-printed-or-not 3d-printed-or-not: An Image Dataset of 3D-printed Prototypes This dataset is a collection of images that are particularly relevant to engineering and design, consisting of two categories: 3D-printed prototypes, and non-3D-printed prototypes This data was collected through a hybrid approach that entailed both web scraping and direct collection from engineering labs and workspaces at Penn State University. The initial data was then augmented using several data augmentation techniques… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/3d-printed-or-not.imageimage-classification10K<n<100K1 likes105 downloads4y agoHugging Face06Francesco /printed-circuit-board Dataset Card for printed-circuit-board ** The original COCO dataset is stored at dataset.tar.gz** Dataset Summary printed-circuit-board Supported Tasks and Leaderboards object-detection: The dataset can be used to train a model for Object Detection. Languages English Dataset Structure Data Instances A data point comprises an image and its object annotations. { 'image_id': 15, 'image': <PIL.JpegImagePlugin.JpegImageFile… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/printed-circuit-board.imageobject-detectionn<1K0 likes41 downloads3y agoHugging Face07Akshit03 /AkshitMajorProjectMIR1_strawberry_printed_session3_20260721_132013This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Akshit03/AkshitMajorProjectMIR1_strawberry_printed_session3_20260721_132013.tabularrobotics1K<n<10K0 likes31 downloads2mo agoHugging Face08arobin79 /bangla-ocr-validation_data_printed Bangla OCR Validation Dataset (Printed + Scanned) 📌 Description This dataset is a Bangla OCR validation dataset containing a mix of printed document images and their corresponding text annotations. It is designed to evaluate OCR and vision-language models on both clean digital text and scanned document images. 📊 Dataset Composition 1507 line-level images with text annotations 50 full-page document images with text Data includes: Printed/typed Bangla text… See the full description on the dataset page: https://huggingface.co/datasets/arobin79/bangla-ocr-validation_data_printed.imageimage-to-text1K<n<10K1 likes30 downloads5mo agoHugging Face09Akshit03 /AkshitMajorProjectMIR1_strawberry_printed_session2_20260721_125019This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Akshit03/AkshitMajorProjectMIR1_strawberry_printed_session2_20260721_125019.tabularrobotics1K<n<10K0 likes30 downloads2mo agoHugging Face10Akshit03 /AkshitMajorProjectMIR1_strawberry_printed_session1_20260721_123935This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Akshit03/AkshitMajorProjectMIR1_strawberry_printed_session1_20260721_123935.tabularrobotics1K<n<10K0 likes28 downloads2mo agoHugging Face11thekamilya /kazakh-printed-dataset Kazakh Printed Dataset for OCR task Data Lineage This dataset was synthetically generated using issai/kazparc as the base. Since Kazakh OCR data is scarce, I developed a pipeline to transform digital Kazakh text into a printed-style dataset. Generation Process Source: Text samples were extracted from issai/kazparc. Augmentation & Stylization: Random Background Color: Simulates different lighting conditions by alternating between… See the full description on the dataset page: https://huggingface.co/datasets/thekamilya/kazakh-printed-dataset.image1K<n<10K0 likes23 downloads5mo agoHugging Face12deepcopy /DonkeySmall-OCR-Cyrillic-Printed-8image1M<n<10M1 likes9 downloads1y agoHugging Face13AKKI-AFK /testset_printed_splitsimage10K<n<100K0 likes7 downloads1y agoHugging Face14electricsheepeurope /europe-owid-production-printed-books-half-century Production Printed Books Half Century | Europe (Our World in Data) 🇪🇺 77 observations · 11 Europe countries · 1475–1775 · Repackaged by Electric Sheep Europe TL;DR This dataset contains 77 observations of Production Printed Books Half Century data across 11 Europe countries, spanning 1475–1775. About the source Source: Our World in Data Publisher: Our World in Data License: cc-by-4.0 Topic: Production Printed Books Half Century… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-production-printed-books-half-century.tabulartabular-classificationn<1K0 likes7 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.