datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-manuscript-dataset
Synthetic Manuscript Dataset
Synthetic historical manuscript folios generated using an automated Python pipeline.
Scripts
The dataset contains three script configurations:
Devanagari
Modi
Sharada
Each script contains 100 synthetic manuscript folios.
Dataset Splits
Split
Samples per Script
Train
85
Validation
10
Test
5
Total
100
Across all three scripts, the dataset contains:
300 manuscript images
300 corresponding Markdown… See the full description on the dataset page: https://huggingface.co/datasets/ved1245/synthetic-manuscript-dataset.synthetic-manuscriptssynthetic-manuscript-generatorMANUS-HaGRID
MANUS-HaGRID: HaGRID-derived Multimodal Annotated Naturalistic Hand Understanding Dataset
MANUS-HaGRID is the HaGRID/HaGRIDv2-derived subset of the Multimodal Annotated Naturalistic Hand Understanding (MANUS) dataset family. It provides multimodal annotations for naturalistic hand gesture understanding, including RGB images, hand crops, depth maps, 2D bounding boxes, estimated MANO-style hand mesh metadata, and multi-view mesh renderings where available.
This repository contains… See the full description on the dataset page: https://huggingface.co/datasets/QFun/MANUS-HaGRID.Arabic_Manuscript_Collection_Dataset
Arabic Manuscript Collection
Seven Arabic handwritten text recognition (HTR) subsets. Five are converted to one layout
and one label format so they can be trained and evaluated together: 82,561 labelled
images in total. Two are republished closer to their source shape: AMIDDA as upstream
Parquet, and OpenITI-Makhzan as page images with line-level coordinates.
Four of the five converted sources are historical manuscripts. KHATT is modern handwriting
and is included as a separate… See the full description on the dataset page: https://huggingface.co/datasets/TheSeniorTeam/Arabic_Manuscript_Collection_Dataset.yalta_ai_segmonto_manuscript_dataset
YALTAi SegmOnto Manuscript and Early Printed Book Dataset
1,147 page images from manuscripts and early printed books, 9th to 17th century, with bounding-box annotations for zone types drawn from the SegmOnto vocabulary. Created by Thibault Clérice and deposited on Zenodo alongside the paper You Actually Look Twice At it (YALTAi) (Journal of Data Mining and Digital Humanities, 2022), which treats page layout recognition on historical documents as an object detection problem… See the full description on the dataset page: https://huggingface.co/datasets/biglam/yalta_ai_segmonto_manuscript_dataset.synthetic-manuscript-generator
Synthetic Manuscript Generator
Synthetic Indic manuscript folios (paper + palm-leaf backgrounds) for OCR
training. Three scripts are produced as separate subsets/configs:
devanagari — 100 folios (85/10/5)
modi — 100 folios (85/10/5)
sharada — 100 folios (85/10/5)
Layout
Each subset is structured as a Hugging Face imagefolder:
<subset>/
train/
0000.png 0000.md metadata.jsonl
...
validation/
...
test/
...
metadata.jsonl rows look like:… See the full description on the dataset page: https://huggingface.co/datasets/Sampada22/synthetic-manuscript-generator.omar-al-saleh-manuscripts-segments
Omar Al-Saleh Manuscripts — Segments
Line-level segmented images with transcriptions from the Omar Al-Saleh memoir collection (1951–1965), part of the NAKBA NLP 2026: Arabic Manuscript Understanding Shared Task.
Dataset
Split
Images
With text
train
15,969
15,969
test
2,095
2,095
blind_test
2,671
2,671
Each example contains:
image: A cropped line image from a manuscript page (JPG or PNG)
text: The Arabic transcription of that line
filename: Original… See the full description on the dataset page: https://huggingface.co/datasets/U4RASD/omar-al-saleh-manuscripts-segments.indic-historical-manuscripts
Synthetic Indic Manuscript Dataset (Devanagari, Modi, Sharada)
This dataset contains synthetic historical manuscript folios paired with exact Markdown (.md) ground-truth transcriptions.
Subsets and Distribution
Subsets: devanagari, modi, sharada
Standard splits:
train: 85%
validation: 10%
test: 5%
Features & Physical Fidelity
Backgrounds: High-resolution procedural aged paper (pothi) & palm-leaf (talapatra) folios with string hole punch marks… See the full description on the dataset page: https://huggingface.co/datasets/varunbhoyar/indic-historical-manuscripts.manuscript_noisy_labelsmanuscript_noisy_labels_iiifmetaboverse-manuscriptRapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts
Dataset Attribution
The original dataset is available on Kaggle.
This dataset has been curated solely for ease of use within the Hugging Face ecosystem, with no intention of plagiarizing or copying the original work.
Please cite the original authors if you use this dataset.
Citation
@INPROCEEDINGS{8978005,
author={Chamchong, Rapeeporn and Gao, Wei and McDonnell, Mark D.},
booktitle={2019 International Conference on Document Analysis and Recognition (ICDAR)}… See the full description on the dataset page: https://huggingface.co/datasets/fwgpiyawudk/RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts.synthetic-indic-manuscriptsMANUS-DexYCB
MANUS-DexYCB: DexYCB-derived Multimodal Annotated Naturalistic Hand Understanding Dataset
MANUS-DexYCB is the DexYCB-derived subset of the Multimodal Annotated Naturalistic Hand Understanding (MANUS) dataset family. It is released as a source-specific repository because MANUS subsets are governed by different upstream licenses.
This repository contains only the DexYCB-derived MANUS test split. HaGRID/HaGRIDv2-derived data is released separately as MANUS-HaGRID.… See the full description on the dataset page: https://huggingface.co/datasets/QFun/MANUS-DexYCB.synthetic-manuscript-generatorrubenstein-manuscript-catalog
Duke University Rubenstein Library Manuscript Card Catalog Dataset
Dataset Description
This dataset contains approximately 48,000 digitized catalog cards from the Duke University David M. Rubenstein Rare Book & Manuscript Library card catalog. The cards represent manuscript collections held by the library, documenting personal papers, organizational records, and historical manuscripts.
Why This Dataset?
Historical manuscript catalog cards are valuable sources… See the full description on the dataset page: https://huggingface.co/datasets/biglam/rubenstein-manuscript-catalog.Tridis_layout_manuscripts
A Unified Dataset for Codicological Document Layout Analysis
Dataset Description
This repository contains a large-scale, unified dataset for Document Layout Analysis (DLA) in historical manuscripts. It was created by harmonizing three distinct public corpora—e-NDP, CATMuS, and HORAE—which cover a wide range of document types from the 12th to the 17th century (administrative registers, literary manuscripts, printed books, and Books of Hours).
The key feature of this… See the full description on the dataset page: https://huggingface.co/datasets/magistermilitum/Tridis_layout_manuscripts.manuscript_iiif_testarmenian-manuscript-htr
Armenian Manuscript HTR (BnF Arménien 172 and UCLA Armenian MS 72)
This release contains handwritten text recognition ground truth for two Armenian canon law manuscripts: 1,704 transcribed lines with line polygons and baselines across 34 pages. Page images ship for both manuscripts: the BnF manuscript under Gallica terms and the UCLA manuscript by decision of the NOMOS project.
The language and the manuscripts
Armenian is an Indo-European language, attested in… See the full description on the dataset page: https://huggingface.co/datasets/nomikos-project/armenian-manuscript-htr.iiif_manuscripts_label_ge_50synthetic-manuscript-datasetbph_manuscript_collectionThe Bibliotheca Philosophica Hermetica is a library containing a broad collection of books on Hermetic philosophy, magic, and the occult. This dataset is a subset of the books from that library:
Filtered down to books in english
Transcribed into text files
Stored in JSON with a mapping of filename -> content for each page.
These books are from the scanned collection at the Embassy of the Free Mind and are from the early 20th century and before.
greek-manuscript-htr
Byzantine Greek Manuscript HTR (BnF Grec 1360)
This release contains handwritten text recognition ground truth for five pages of Paris, Bibliothèque nationale de France, Grec 1360, a 1351 copy of the Hexabiblos of Constantine Harmenopoulos: 204 transcribed lines with line polygons and baselines, plus the five page images.
The language and the manuscripts
Byzantine or medieval Greek was the principal language of the East Roman state and the Orthodox Church through… See the full description on the dataset page: https://huggingface.co/datasets/nomikos-project/greek-manuscript-htr.syriac-manuscript-htr
East Syriac Manuscript HTR (Alqosh 65)
This release contains handwritten text recognition ground truth for seven pages of Alqosh 65, a nineteenth-century East Syriac manuscript that contains a letter by a jurist-bishop-philosopher on the ten categories of Aristotle: 143 transcribed lines with line polygons and baselines, plus the seven page images. Geometry and text are the project's own work.
The language and the manuscripts
Syriac is the language of various… See the full description on the dataset page: https://huggingface.co/datasets/nomikos-project/syriac-manuscript-htr.MANUS
MANUS Dataset Family
Multimodal Annotated Naturalistic Hand Understanding (MANUS) is a family of source-specific multimodal hand-understanding dataset releases for naturalistic hand analysis, hand-aware image generation, 2D/3D hand understanding, depth estimation, and multi-view hand representation learning.
MANUS does not use a single dataset-wide license. Each source-specific release is governed by its own upstream-compatible license. This landing repository indexes the available… See the full description on the dataset page: https://huggingface.co/datasets/QFun/MANUS.rubenstein-manuscript-catalog-glm-ocr
Document OCR using GLM-OCR
This dataset contains OCR results from images in biglam/rubenstein-manuscript-catalog using GLM-OCR, a compact 0.9B OCR model achieving SOTA performance.
Processing Details
Source Dataset: biglam/rubenstein-manuscript-catalog
Model: zai-org/GLM-OCR
Task: text recognition
Number of Samples: 49,654
Processing Time: 343.2 min
Processing Date: 2026-02-14 15:38 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/rubenstein-manuscript-catalog-glm-ocr.manusctips_data_for_OCR_splittedcxr-10k-reasoning-dataset
🫁 CXR-10K Reasoning Dataset
A dataset of 10,000 chest X-ray images paired with step-by-step clinical reasoning and radiology impression summaries, curated for training and evaluating medical vision-language models like MedGEMMA, LLaVA-Med, and others.
📂 Dataset Structure
This dataset is saved in Arrow format and was built using the Hugging Face datasets library.
Each sample includes:
image: Chest X-ray image (PNG or JPEG)
reasoning: Step-wise radiological reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/Manusinhh/cxr-10k-reasoning-dataset.synthetic-indic-manuscripts
