lightonocr
LightOnOCR-mix-0126
LightOnOCR-mix-0126
LightOnOCR-mix-0126 is a large-scale OCR training dataset built via distillation: a strong vision–language model is prompted to produce naturally ordered full-page transcriptions (Markdown with LaTeX math spans and HTML tables) from rendered document pages. The dataset is designed as supervision for end-to-end OCR / document-understanding models that aim to output clean, human-readable text in a consistent format.
This repository releases the PDFA-derived… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/LightOnOCR-mix-0126.LightOnOCR-bbox-mix-0126
LightOnOCR-bbox-mix-0126
LightOnOCR-bbox-mix-0126 is a large-scale OCR training dataset including layout information built via distillation: a strong vision–language model is prompted to produce naturally ordered full-page transcriptions (Markdown with LaTeX math spans and HTML tables) from rendered document pages. The dataset is designed as supervision for end-to-end OCR / document-understanding models that aim to output clean, human-readable text in a consistent format.
This… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/LightOnOCR-bbox-mix-0126.lightonocr-pubtableslightonocr-pubtables-table-onlysamaritan_hebrew_LightOnOcr
Samaritan Hebrew OCR Dataset
Dataset Summary
The Samaritan Hebrew OCR Dataset is a specialized dataset for fine-tuning OCR models on Samaritan Hebrew manuscripts. This dataset contains 46,860 annotated samples extracted from 1,374 manuscript pages, converted from PAGE-XML format to the LightOnOCR-2 training format.
The dataset includes three types of samples:
Line-level samples: Individual textlines cropped using precise polygon masks (40,219 samples)
Paragraph-level… See the full description on the dataset page: https://huggingface.co/datasets/samaritan-ai/samaritan_hebrew_LightOnOcr.LightOnOCR-bbox-bench
LightOnOCR-bbox-bench
Evaluation benchmark for assessing the ability of vision-language models (VLMs) to localize images within documents using bounding boxes. This dataset was introduced in the paper LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR.
Task Description
Given a document page (PDF), the model must predict bounding boxes around images (figures, charts, photographs, etc.) present in the document. This evaluates the model's… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/LightOnOCR-bbox-bench.
