manuscript
synthetic-manuscript-dataset
Synthetic Manuscript Dataset
Synthetic historical manuscript folios generated using an automated Python pipeline.
Scripts
The dataset contains three script configurations:
Devanagari
Modi
Sharada
Each script contains 100 synthetic manuscript folios.
Dataset Splits
Split
Samples per Script
Train
85
Validation
10
Test
5
Total
100
Across all three scripts, the dataset contains:
300 manuscript images
300 corresponding Markdown… See the full description on the dataset page: https://huggingface.co/datasets/ved1245/synthetic-manuscript-dataset.synthetic-manuscriptssynthetic-manuscript-generatorArabic_Manuscript_Collection_Dataset
Arabic Manuscript Collection
Seven Arabic handwritten text recognition (HTR) subsets. Five are converted to one layout
and one label format so they can be trained and evaluated together: 82,561 labelled
images in total. Two are republished closer to their source shape: AMIDDA as upstream
Parquet, and OpenITI-Makhzan as page images with line-level coordinates.
Four of the five converted sources are historical manuscripts. KHATT is modern handwriting
and is included as a separate… See the full description on the dataset page: https://huggingface.co/datasets/TheSeniorTeam/Arabic_Manuscript_Collection_Dataset.yalta_ai_segmonto_manuscript_dataset
YALTAi SegmOnto Manuscript and Early Printed Book Dataset
1,147 page images from manuscripts and early printed books, 9th to 17th century, with bounding-box annotations for zone types drawn from the SegmOnto vocabulary. Created by Thibault Clérice and deposited on Zenodo alongside the paper You Actually Look Twice At it (YALTAi) (Journal of Data Mining and Digital Humanities, 2022), which treats page layout recognition on historical documents as an object detection problem… See the full description on the dataset page: https://huggingface.co/datasets/biglam/yalta_ai_segmonto_manuscript_dataset.synthetic-manuscript-generator
Synthetic Manuscript Generator
Synthetic Indic manuscript folios (paper + palm-leaf backgrounds) for OCR
training. Three scripts are produced as separate subsets/configs:
devanagari — 100 folios (85/10/5)
modi — 100 folios (85/10/5)
sharada — 100 folios (85/10/5)
Layout
Each subset is structured as a Hugging Face imagefolder:
<subset>/
train/
0000.png 0000.md metadata.jsonl
...
validation/
...
test/
...
metadata.jsonl rows look like:… See the full description on the dataset page: https://huggingface.co/datasets/Sampada22/synthetic-manuscript-generator.
