datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
britannica-illustrated-pages
Britannica Illustrated Pages
115,293 illustrated pages from scanned volumes of the Encyclopaedia Britannica, 1st edition
(1768–71) to 14th (1929), selected by a page classifier from 975,345 pages in 1,160 volumes
(838 Internet Archive items). A second config carries the classifier
score, OCR word count and provenance for every one of the 975,345 pages.
Two things the scan showed:
82% of the illustrated pages are text pages (≥100 OCR words) — figures, diagrams and
engravings set… See the full description on the dataset page: https://huggingface.co/datasets/biglam/britannica-illustrated-pages.od-syn-page-annotations-com
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis, visual document understanding, OCR fine-tuning, and related tasks specifically for Dhivehi script.
Note: this version image are compressed.
Raw version 📁 Repository: Hugging Face Datasets
📋 Dataset Summary
Total Examples: ~58… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations-com.tibetan-page-orientation-classifier-dataset
Tibetan Page Orientation Dataset
Covers 7 Tibetan script families: Danyig, Druma, Gyuyig, Multi-Scripts, Pedri, Tsugdri, Uchen.
Dataset composition
Each manuscript page appears twice: once as the original scan (non_flipped) and once rotated 180° (flipped). The model's task is to distinguish these two orientations.
Scripts are balanced — each of the 7 script families contributes the same number of pages (downsampled to the smallest family).
Script (script)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-page-orientation-classifier-dataset.autotrain-data-page7
AutoTrain Dataset for project: page7
Dataset Description
This dataset has been automatically processed by AutoTrain for project page7.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<241x411 RGB PIL image>",
"target": 6
},
{
"image": "<209x293 RGB PIL image>",
"target": 1
}]
Dataset Fields
The… See the full description on the dataset page: https://huggingface.co/datasets/rdmpage/autotrain-data-page7.vibe-landing-page-arena
Vibe Landing Page Arena
A large-scale human preference dataset for evaluating AI-generated landing page design quality. 36,000 pairwise judgments from 3,492 annotators comparing landing pages generated by Claude Code, Cursor, Lovable, and Replit across 100 prompts and 4 design dimensions.
Overview
Metric
Value
Total judgments
36,000
Unique annotators
3,492
Prompts
100
Business categories
97
Design tones
82
Tools compared
4 (Claude Code, Cursor… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/vibe-landing-page-arena.autotrain-data-pagex
AutoTrain Dataset for project: pagex
Dataset Description
This dataset has been automatically processed by AutoTrain for project pagex.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<235x313 RGB PIL image>",
"target": 1
},
{
"image": "<235x313 RGB PIL image>",
"target": 0
}]
Dataset Fields
The… See the full description on the dataset page: https://huggingface.co/datasets/rdmpage/autotrain-data-pagex.8-class-tibetan-page-classification-dataset
8-Class Tibetan Page Classification
Page-level classification of BDRC manuscript / print images into 8 classes: the
six script categories of BDRC/6-class-tibetan-script-classification-dataset
plus two page-type classes — blank and nonplaintext — so a downstream
OCR pipeline can route pages (skip blanks, handle tables/illustrations/scores separately).
Trained classifier: BDRC/8-class-tibetan-page-classifier.
Classes
Class
Description
danyig_pedri
Danyig… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/8-class-tibetan-page-classification-dataset.incunabula-pages
Dataset Card for Incunabula Bibles Illustration Pages
This is a collection of pages from early modern Bibles categorized into a) containing illustrations, b) text-only.
It was used to train the Incunabula Page Classifier.
Dataset Structure
data/
├── train/
│ ├── illustrated/ # Pages containing illustrations
│ └── text_only/ # Pages with only text
└── val/
├── illustrated/ # Validation pages with illustrations
└── text_only/ # Validation… See the full description on the dataset page: https://huggingface.co/datasets/zghem/incunabula-pages.
