CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NealCaren /newspaper-pagesimage0 likes5.2k downloads2mo agoHugging Face02biglam /britannica-illustrated-pages Britannica Illustrated Pages 115,293 illustrated pages from scanned volumes of the Encyclopaedia Britannica, 1st edition (1768–71) to 14th (1929), selected by a page classifier from 975,345 pages in 1,160 volumes (838 Internet Archive items). A second config carries the classifier score, OCR word count and provenance for every one of the 975,345 pages. Two things the scan showed: 82% of the illustrated pages are text pages (≥100 OCR words) — figures, diagrams and engravings set… See the full description on the dataset page: https://huggingface.co/datasets/biglam/britannica-illustrated-pages.imageimage-classification1M<n<10M49 likes3.3k downloads1mo agoHugging Face03nielsr /paper-page-assetsdocumentn<1K1 likes3k downloads1y agoHugging Face04pixparse /docvqa-single-page-questions Dataset Card for DocVQA Dataset Dataset Summary DocVQA dataset is a document dataset introduced in Mathew et al. (2021) consisting of 50,000 questions defined on 12,000+ document images. Please visit the challenge page (https://rrc.cvc.uab.es/?ch=17) and paper (https://arxiv.org/abs/2007.00398) for further information. Usage This dataset can be used with current releases of Hugging Face datasets library. Here is an example using a custom collator to bundle… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/docvqa-single-page-questions.imagequestion-answering10K<n<100K11 likes2.8k downloads2y agoHugging Face05Reza2kn /persian-handwriting-pages-3.69m Persian Handwriting Pages 3.69M 3,690,000 deterministic, densely composed Persian handwriting pages. This expansion uses new random seeds and is complementary to Reza2kn/persian-handwriting-pages-369k, not a repetition of its rendered pages. The public viewer intentionally exposes exactly two columns: image and label. Pages are uploaded as verified Parquet shards and deleted locally after remote-size verification. Source handwriting Word images originate from Taha… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-3.69m.imageimage-to-text1M<n<10M4 likes2.6k downloads2mo agoHugging Face06RoboCOIN /pageAssetsimagen<1K0 likes2.2k downloads6d agoHugging Face07VLM2Vec /MMLongBench-page-fixedimage1K<n<10K0 likes2k downloads11mo agoHugging Face08VLM2Vec /ViDoSeek-page-fixedimage1K<n<10K0 likes2k downloads11mo agoHugging Face09huanngzh /page-assets3dn<1K0 likes2k downloads5d agoHugging Face10Reza2kn /persian-handwriting-pages-369k Persian Handwriting Pages 369K Full-page Persian handwriting compositions on scanned paper backgrounds. Each row deliberately has only two fields: image: the composed full-page image label: its complete line-separated Persian transcription, ordered from top to bottom The pages are composed from labeled real handwriting crops with page-level ink normalization, controlled RTL layout variation, collision prevention, and exact transcription provenance. The release contains 369,000… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-369k.imageimage-to-text100K<n<1M4 likes1.1k downloads2mo agoHugging Face11Joinn /Pageimagen<1K0 likes934 downloads1mo agoHugging Face12nojiyoon /pagoda-text-and-image-dataset Dataset Card for "pagoda-text-and-image-dataset" More Information needed imagen<1K1 likes527 downloads3y agoHugging Face13alakxender /od-syn-page-annotations-com 📦 Dhivehi Synthetic Document Layout + Textline Dataset This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis, visual document understanding, OCR fine-tuning, and related tasks specifically for Dhivehi script. Note: this version image are compressed. Raw version 📁 Repository: Hugging Face Datasets 📋 Dataset Summary Total Examples: ~58… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations-com.imageimage-classification10K<n<100K0 likes464 downloads1y agoHugging Face14alakxender /od-syn-page-annotations 📦 Dhivehi Synthetic Document Layout + Textline Dataset This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis , visual document understanding , OCR fine-tuning, and related tasks specifically for Dhivehi script. 📋 Dataset Summary Total Examples: ~58,738 Image Content: Synthetic Dhivehi documents generated to simulate real-world layouts… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations.imagetext-classification10K<n<100K0 likes399 downloads1y agoHugging Face15Sudith /svlm-financial-pagesimage1K<n<10K0 likes333 downloads2mo agoHugging Face16huggingface /task-page-imagesaudion<1K0 likes290 downloads5y agoHugging Face17mriosqu /landing_pages_04_dataset Dataset Card for "landing_pages_04_dataset" More Information needed imagen<1K1 likes278 downloads3y agoHugging Face18PAGF /DET-COMPASS Superpowering Open-Vocabulary Object Detectors for X-ray Vision ICCV 2025 Pablo Garcia-Fernandez, Lorenzo Vaquero, Mingxuan Liu, Feng Xue, Daniel Cores, Nicu Sebe, Manuel Mucientes, Elisa Ricci DET-COMPASS This is the official repository of Superpowering Open-Vocabulary Object Detectors for X-ray Vision (ICCV'25) Dataset Summary Object detection in security X-ray scans has advanced significantly in recent years. However, evaluating Open-vocabulary Object… See the full description on the dataset page: https://huggingface.co/datasets/PAGF/DET-COMPASS.imageobject-detection1K<n<10K2 likes268 downloads1y agoHugging Face19BDRC /tibetan-page-orientation-classifier-dataset Tibetan Page Orientation Dataset Covers 7 Tibetan script families: Danyig, Druma, Gyuyig, Multi-Scripts, Pedri, Tsugdri, Uchen. Dataset composition Each manuscript page appears twice: once as the original scan (non_flipped) and once rotated 180° (flipped). The model's task is to distinguish these two orientations. Scripts are balanced — each of the 7 script families contributes the same number of pages (downsampled to the smallest family). Script (script)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-page-orientation-classifier-dataset.imageimage-classification1K<n<10K0 likes263 downloads3mo agoHugging Face20rdmpage /autotrain-data-page7 AutoTrain Dataset for project: page7 Dataset Description This dataset has been automatically processed by AutoTrain for project page7. Languages The BCP-47 code for the dataset's language is unk. Dataset Structure Data Instances A sample from this dataset looks as follows: [ { "image": "<241x411 RGB PIL image>", "target": 6 }, { "image": "<209x293 RGB PIL image>", "target": 1 }] Dataset Fields The… See the full description on the dataset page: https://huggingface.co/datasets/rdmpage/autotrain-data-page7.imageimage-classification0 likes256 downloads3y agoHugging Face21dfs-team /mathglyph-pages MathGlyph Pages 1k With Detector Boxes Synthetic mixed handwritten math pages prepared for the DFS/Rukopys detector pretraining pipeline. Links Generator repo: reirei-00/mathglyph_pages Format train/images/: train page images. validation/images/: validation page images. annotations/instance_train.json: COCO-style detector annotations for train. annotations/instance_val.json: COCO-style detector annotations for validation. train/metadata.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/dfs-team/mathglyph-pages.imageobject-detection1K<n<10K0 likes240 downloads3mo agoHugging Face22WenyiWU0111 /structagent-page-assetsimage1K<n<10K0 likes226 downloads3mo agoHugging Face23medieval-data /catmus-caroline-pageimage1K<n<10K0 likes186 downloads2y agoHugging Face24nojiyoon /pagoda-text-and-image-dataset-small Dataset Card for "pagoda-text-and-image-dataset-small" More Information needed imagen<1K0 likes171 downloads3y agoHugging Face25EduardoPacheco /Fox-Page-En Fox-Page-En Dataset comprising the English subset of Fox dataset for the pdf pages. This is just a convenient way to access this subset from the original dataset repo imagen<1K0 likes154 downloads1y agoHugging Face26DoctorSlimm /mozart-api-demo-pages Dataset Card for Dataset Name Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DoctorSlimm/mozart-api-demo-pages.imagen<1K0 likes140 downloads3y agoHugging Face27TheRealOKAI /muharaf-public-pages Muharaf-public images This dataset contains 1,218 full page images of Arabic handwriting and the corresponding text. The images are Manuscripts from the 19th to 21st century. See the official code, paper, and zenodo archive below. Their work has been accepted to NeurIPS 2024. How to use from datasets import load_dataset import matplotlib.pyplot as plt # Load your dataset in streaming mode to be loaded faster #… See the full description on the dataset page: https://huggingface.co/datasets/TheRealOKAI/muharaf-public-pages.image1K<n<10K7 likes135 downloads10mo agoHugging Face28FatimahEmadEldin /Gutenberg-Arabic-OCR-HTML-Pages Gutenberg Arabic HTML-Page Dataset 📖 Dataset Description The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language. The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.image1K<n<10K3 likes129 downloads1y agoHugging Face29emanuelevivoli /comix_v0_tiny_pages Comic Books Tiny Dataset v0 - Pages (Testing) Small test dataset of comic book pages for rapid development and testing. ⚠️ This is a TINY dataset for testing only. For production, use comix_v0_pages. What's Included Each page has: {page_id}.jpg - Page image {page_id}.json - Metadata (detections, captions, page class) {page_id}.seg.npz - Segmentation masks (SAMv2) Quick Start from datasets import load_dataset import numpy as np # Load tiny pages dataset pages… See the full description on the dataset page: https://huggingface.co/datasets/emanuelevivoli/comix_v0_tiny_pages.imageimage-to-text1K<n<10K0 likes105 downloads10mo agoHugging Face30datapointai /vibe-landing-page-arenagated Vibe Landing Page Arena A large-scale human preference dataset for evaluating AI-generated landing page design quality. 36,000 pairwise judgments from 3,492 annotators comparing landing pages generated by Claude Code, Cursor, Lovable, and Replit across 100 prompts and 4 design dimensions. Overview Metric Value Total judgments 36,000 Unique annotators 3,492 Prompts 100 Business categories 97 Design tones 82 Tools compared 4 (Claude Code, Cursor… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/vibe-landing-page-arena.imageimage-classification1K<n<10K2 likes93 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.