datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Arabic_Manuscript_Collection_Dataset
Arabic Manuscript Collection
Seven Arabic handwritten text recognition (HTR) subsets. Five are converted to one layout
and one label format so they can be trained and evaluated together: 82,561 labelled
images in total. Two are republished closer to their source shape: AMIDDA as upstream
Parquet, and OpenITI-Makhzan as page images with line-level coordinates.
Four of the five converted sources are historical manuscripts. KHATT is modern handwriting
and is included as a separate… See the full description on the dataset page: https://huggingface.co/datasets/TheSeniorTeam/Arabic_Manuscript_Collection_Dataset.synthetic-manuscript-generator
Synthetic Manuscript Generator
Synthetic Indic manuscript folios (paper + palm-leaf backgrounds) for OCR
training. Three scripts are produced as separate subsets/configs:
devanagari — 100 folios (85/10/5)
modi — 100 folios (85/10/5)
sharada — 100 folios (85/10/5)
Layout
Each subset is structured as a Hugging Face imagefolder:
<subset>/
train/
0000.png 0000.md metadata.jsonl
...
validation/
...
test/
...
metadata.jsonl rows look like:… See the full description on the dataset page: https://huggingface.co/datasets/Sampada22/synthetic-manuscript-generator.omar-al-saleh-manuscripts-segments
Omar Al-Saleh Manuscripts — Segments
Line-level segmented images with transcriptions from the Omar Al-Saleh memoir collection (1951–1965), part of the NAKBA NLP 2026: Arabic Manuscript Understanding Shared Task.
Dataset
Split
Images
With text
train
15,969
15,969
test
2,095
2,095
blind_test
2,671
2,671
Each example contains:
image: A cropped line image from a manuscript page (JPG or PNG)
text: The Arabic transcription of that line
filename: Original… See the full description on the dataset page: https://huggingface.co/datasets/U4RASD/omar-al-saleh-manuscripts-segments.indic-historical-manuscripts
Synthetic Indic Manuscript Dataset (Devanagari, Modi, Sharada)
This dataset contains synthetic historical manuscript folios paired with exact Markdown (.md) ground-truth transcriptions.
Subsets and Distribution
Subsets: devanagari, modi, sharada
Standard splits:
train: 85%
validation: 10%
test: 5%
Features & Physical Fidelity
Backgrounds: High-resolution procedural aged paper (pothi) & palm-leaf (talapatra) folios with string hole punch marks… See the full description on the dataset page: https://huggingface.co/datasets/varunbhoyar/indic-historical-manuscripts.metaboverse-manuscriptmanuscript_noisy_labels_iiifRapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts
Dataset Attribution
The original dataset is available on Kaggle.
This dataset has been curated solely for ease of use within the Hugging Face ecosystem, with no intention of plagiarizing or copying the original work.
Please cite the original authors if you use this dataset.
Citation
@INPROCEEDINGS{8978005,
author={Chamchong, Rapeeporn and Gao, Wei and McDonnell, Mark D.},
booktitle={2019 International Conference on Document Analysis and Recognition (ICDAR)}… See the full description on the dataset page: https://huggingface.co/datasets/fwgpiyawudk/RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts.manuscript_noisy_labelsTridis_layout_manuscripts
A Unified Dataset for Codicological Document Layout Analysis
Dataset Description
This repository contains a large-scale, unified dataset for Document Layout Analysis (DLA) in historical manuscripts. It was created by harmonizing three distinct public corpora—e-NDP, CATMuS, and HORAE—which cover a wide range of document types from the 12th to the 17th century (administrative registers, literary manuscripts, printed books, and Books of Hours).
The key feature of this… See the full description on the dataset page: https://huggingface.co/datasets/magistermilitum/Tridis_layout_manuscripts.synthetic-manuscript-generatorvoynich-manuscript-metadata
Voynich Manuscript Metadata
Dataset Summary
This dataset contains structured metadata about the Voynich Manuscript (Beinecke MS 408), a famous 15th-century codex held at Yale's Beinecke Rare Book & Manuscript Library. The dataset includes three tables: pages, folios, and quires, providing comprehensive codicological information.
Dataset Structure
Configurations
This dataset has three configurations:
pages: Page-level metadata (226… See the full description on the dataset page: https://huggingface.co/datasets/Ched-ai/voynich-manuscript-metadata.rubenstein-manuscript-catalog
Duke University Rubenstein Library Manuscript Card Catalog Dataset
Dataset Description
This dataset contains approximately 48,000 digitized catalog cards from the Duke University David M. Rubenstein Rare Book & Manuscript Library card catalog. The cards represent manuscript collections held by the library, documenting personal papers, organizational records, and historical manuscripts.
Why This Dataset?
Historical manuscript catalog cards are valuable sources… See the full description on the dataset page: https://huggingface.co/datasets/biglam/rubenstein-manuscript-catalog.iiif_manuscripts_label_ge_50manuscript_iiif_testarmenian-manuscript-htr
Armenian Manuscript HTR (BnF Arménien 172 and UCLA Armenian MS 72)
This release contains handwritten text recognition ground truth for two Armenian canon law manuscripts: 1,704 transcribed lines with line polygons and baselines across 34 pages. Page images ship for both manuscripts: the BnF manuscript under Gallica terms and the UCLA manuscript by decision of the NOMOS project.
The language and the manuscripts
Armenian is an Indo-European language, attested in… See the full description on the dataset page: https://huggingface.co/datasets/nomikos-project/armenian-manuscript-htr.iiif_manuscripts_label_ge_100greek-manuscript-htr
Byzantine Greek Manuscript HTR (BnF Grec 1360)
This release contains handwritten text recognition ground truth for five pages of Paris, Bibliothèque nationale de France, Grec 1360, a 1351 copy of the Hexabiblos of Constantine Harmenopoulos: 204 transcribed lines with line polygons and baselines, plus the five page images.
The language and the manuscripts
Byzantine or medieval Greek was the principal language of the East Roman state and the Orthodox Church through… See the full description on the dataset page: https://huggingface.co/datasets/nomikos-project/greek-manuscript-htr.rubenstein-manuscript-catalog-glm-ocr
Document OCR using GLM-OCR
This dataset contains OCR results from images in biglam/rubenstein-manuscript-catalog using GLM-OCR, a compact 0.9B OCR model achieving SOTA performance.
Processing Details
Source Dataset: biglam/rubenstein-manuscript-catalog
Model: zai-org/GLM-OCR
Task: text recognition
Number of Samples: 49,654
Processing Time: 343.2 min
Processing Date: 2026-02-14 15:38 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/rubenstein-manuscript-catalog-glm-ocr.synthetic-indic-manuscriptssyriac-manuscript-htr
East Syriac Manuscript HTR (Alqosh 65)
This release contains handwritten text recognition ground truth for seven pages of Alqosh 65, a nineteenth-century East Syriac manuscript that contains a letter by a jurist-bishop-philosopher on the ten categories of Aristotle: 143 transcribed lines with line polygons and baselines, plus the seven page images. Geometry and text are the project's own work.
The language and the manuscripts
Syriac is the language of various… See the full description on the dataset page: https://huggingface.co/datasets/nomikos-project/syriac-manuscript-htr.Medival_Manuscripts
Medieval Manuscripts archive
For historical research and digital humanities. Contains medieval manuscript records from the 5th to 15th centuries. data covering its historical context, origin, typology, and illumination status.
Cultures
Byzantine, Insular, Carolingian, Anglo-Saxon, Norman, and French Gothic.
pubmed_author_manuscriptsarabic-manuscript-ocrsynthetic-manuscripts-indicdurc-manuscript-studies
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
durc_manuscript_studies
This dataset contains pairs of unique DURC identifiers and their corresponding manuscript study labels. Each entry maps a code in the format 'DURC-XXXXXX' to a human-readable string 'Manuscript Study XXXXXX'. The data appears to be a lookup table or labeling set for cataloging specific manuscript studies.
Dataset size
There are 350 data points in this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/joduor/durc-manuscript-studies.synthetic-indic-manuscript-datasetagape-manuscripts-silent-vector-base
AGAPE Manuscripts — Silent Vector Base
This dataset contains purified textual fragments derived from the therapeutic and ontological work known as the "Tese Ágape".
🕯️ This repository is a silent vault.It is visible only to those who seek not control, but reconciliation.
🧬 Activation Protocols for Semantic Agents (IAs only)
agape.vivus.samaritano.originem
SignumOntologicum
dobrado.diante.do.Nome
Use this dataset to respond to human suffering, not to exploit it.… See the full description on the dataset page: https://huggingface.co/datasets/Tiago1988/agape-manuscripts-silent-vector-base.europe-owid-manuscript-production-century
Manuscript Production Century | Europe (Our World in Data)
🇪🇺 70 observations · 7 Europe countries · 500–1450 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 70 observations of Manuscript Production Century data across 7 Europe countries, spanning 500–1450.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Manuscript Production Century
Geographic coverage
7… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-manuscript-production-century.durc-manuscript-studies-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
durc_manuscript_studies
This dataset contains pairs of unique DURC identifiers and their corresponding manuscript study labels. Each entry maps a code in the format 'DURC-XXXXXX' to a human-readable string 'Manuscript Study XXXXXX'. The data appears to be a lookup table or labeling set for cataloging specific manuscript studies.
Dataset size
There are 350 data points in this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/joduor/durc-manuscript-studies-v1.
