datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
newspaper-pagesbritannica-illustrated-pages
Britannica Illustrated Pages
115,293 illustrated pages from scanned volumes of the Encyclopaedia Britannica, 1st edition
(1768–71) to 14th (1929), selected by a page classifier from 975,345 pages in 1,160 volumes
(838 Internet Archive items). A second config carries the classifier
score, OCR word count and provenance for every one of the 975,345 pages.
Two things the scan showed:
82% of the illustrated pages are text pages (≥100 OCR words) — figures, diagrams and
engravings set… See the full description on the dataset page: https://huggingface.co/datasets/biglam/britannica-illustrated-pages.paper-page-assetsdocvqa-single-page-questions
Dataset Card for DocVQA Dataset
Dataset Summary
DocVQA dataset is a document dataset introduced in Mathew et al. (2021) consisting of 50,000 questions defined on 12,000+ document images.
Please visit the challenge page (https://rrc.cvc.uab.es/?ch=17) and paper (https://arxiv.org/abs/2007.00398) for further information.
Usage
This dataset can be used with current releases of Hugging Face datasets library.
Here is an example using a custom collator to bundle… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/docvqa-single-page-questions.persian-handwriting-pages-3.69m
Persian Handwriting Pages 3.69M
3,690,000 deterministic, densely composed Persian handwriting pages.
This expansion uses new random seeds and is complementary to
Reza2kn/persian-handwriting-pages-369k,
not a repetition of its rendered pages.
The public viewer intentionally exposes exactly two columns: image and label.
Pages are uploaded as verified Parquet shards and deleted locally after remote-size verification.
Source handwriting
Word images originate from Taha… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-3.69m.pageAssetsMMLongBench-page-fixedViDoSeek-page-fixedpage-assetspersian-handwriting-pages-369k
Persian Handwriting Pages 369K
Full-page Persian handwriting compositions on scanned paper backgrounds.
Each row deliberately has only two fields:
image: the composed full-page image
label: its complete line-separated Persian transcription, ordered from top to bottom
The pages are composed from labeled real handwriting crops with page-level ink normalization,
controlled RTL layout variation, collision prevention, and exact transcription provenance.
The release contains 369,000… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-369k.Pagepagoda-text-and-image-dataset
Dataset Card for "pagoda-text-and-image-dataset"
More Information needed
od-syn-page-annotations-com
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis, visual document understanding, OCR fine-tuning, and related tasks specifically for Dhivehi script.
Note: this version image are compressed.
Raw version 📁 Repository: Hugging Face Datasets
📋 Dataset Summary
Total Examples: ~58… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations-com.od-syn-page-annotations
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis , visual document understanding , OCR fine-tuning, and related tasks specifically for Dhivehi script.
📋 Dataset Summary
Total Examples: ~58,738
Image Content: Synthetic Dhivehi documents generated to simulate real-world layouts… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations.svlm-financial-pagestask-page-imageslanding_pages_04_dataset
Dataset Card for "landing_pages_04_dataset"
More Information needed
DET-COMPASS
Superpowering Open-Vocabulary Object Detectors for X-ray Vision
ICCV 2025
Pablo Garcia-Fernandez,
Lorenzo Vaquero,
Mingxuan Liu,
Feng Xue,
Daniel Cores,
Nicu Sebe,
Manuel Mucientes,
Elisa Ricci
DET-COMPASS
This is the official repository of Superpowering Open-Vocabulary Object Detectors for X-ray Vision (ICCV'25)
Dataset Summary
Object detection in security X-ray scans has advanced significantly in recent years. However, evaluating Open-vocabulary Object… See the full description on the dataset page: https://huggingface.co/datasets/PAGF/DET-COMPASS.tibetan-page-orientation-classifier-dataset
Tibetan Page Orientation Dataset
Covers 7 Tibetan script families: Danyig, Druma, Gyuyig, Multi-Scripts, Pedri, Tsugdri, Uchen.
Dataset composition
Each manuscript page appears twice: once as the original scan (non_flipped) and once rotated 180° (flipped). The model's task is to distinguish these two orientations.
Scripts are balanced — each of the 7 script families contributes the same number of pages (downsampled to the smallest family).
Script (script)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-page-orientation-classifier-dataset.autotrain-data-page7
AutoTrain Dataset for project: page7
Dataset Description
This dataset has been automatically processed by AutoTrain for project page7.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<241x411 RGB PIL image>",
"target": 6
},
{
"image": "<209x293 RGB PIL image>",
"target": 1
}]
Dataset Fields
The… See the full description on the dataset page: https://huggingface.co/datasets/rdmpage/autotrain-data-page7.mathglyph-pages
MathGlyph Pages 1k With Detector Boxes
Synthetic mixed handwritten math pages prepared for the DFS/Rukopys detector pretraining pipeline.
Links
Generator repo: reirei-00/mathglyph_pages
Format
train/images/: train page images.
validation/images/: validation page images.
annotations/instance_train.json: COCO-style detector annotations for train.
annotations/instance_val.json: COCO-style detector annotations for validation.
train/metadata.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/dfs-team/mathglyph-pages.structagent-page-assetscatmus-caroline-pagepagoda-text-and-image-dataset-small
Dataset Card for "pagoda-text-and-image-dataset-small"
More Information needed
Fox-Page-En
Fox-Page-En
Dataset comprising the English subset of Fox dataset for the pdf pages. This is just a convenient way to access this subset from the original dataset repo
mozart-api-demo-pages
Dataset Card for Dataset Name
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DoctorSlimm/mozart-api-demo-pages.muharaf-public-pages
Muharaf-public images
This dataset contains 1,218 full page images of Arabic handwriting and the corresponding text. The images are Manuscripts from the 19th to 21st century.
See the official code, paper, and zenodo archive below. Their work has been accepted to NeurIPS 2024.
How to use
from datasets import load_dataset
import matplotlib.pyplot as plt
# Load your dataset in streaming mode to be loaded faster
#… See the full description on the dataset page: https://huggingface.co/datasets/TheRealOKAI/muharaf-public-pages.Gutenberg-Arabic-OCR-HTML-Pages
Gutenberg Arabic HTML-Page Dataset
📖 Dataset Description
The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language.
The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.comix_v0_tiny_pages
Comic Books Tiny Dataset v0 - Pages (Testing)
Small test dataset of comic book pages for rapid development and testing.
⚠️ This is a TINY dataset for testing only. For production, use comix_v0_pages.
What's Included
Each page has:
{page_id}.jpg - Page image
{page_id}.json - Metadata (detections, captions, page class)
{page_id}.seg.npz - Segmentation masks (SAMv2)
Quick Start
from datasets import load_dataset
import numpy as np
# Load tiny pages dataset
pages… See the full description on the dataset page: https://huggingface.co/datasets/emanuelevivoli/comix_v0_tiny_pages.vibe-landing-page-arena
Vibe Landing Page Arena
A large-scale human preference dataset for evaluating AI-generated landing page design quality. 36,000 pairwise judgments from 3,492 annotators comparing landing pages generated by Claude Code, Cursor, Lovable, and Replit across 100 prompts and 4 design dimensions.
Overview
Metric
Value
Total judgments
36,000
Unique annotators
3,492
Prompts
100
Business categories
97
Design tones
82
Tools compared
4 (Claude Code, Cursor… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/vibe-landing-page-arena.
