datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LightOnOCR-mix-0126
LightOnOCR-mix-0126
LightOnOCR-mix-0126 is a large-scale OCR training dataset built via distillation: a strong vision–language model is prompted to produce naturally ordered full-page transcriptions (Markdown with LaTeX math spans and HTML tables) from rendered document pages. The dataset is designed as supervision for end-to-end OCR / document-understanding models that aim to output clean, human-readable text in a consistent format.
This repository releases the PDFA-derived… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/LightOnOCR-mix-0126.LightOnOCR-bbox-mix-0126
LightOnOCR-bbox-mix-0126
LightOnOCR-bbox-mix-0126 is a large-scale OCR training dataset including layout information built via distillation: a strong vision–language model is prompted to produce naturally ordered full-page transcriptions (Markdown with LaTeX math spans and HTML tables) from rendered document pages. The dataset is designed as supervision for end-to-end OCR / document-understanding models that aim to output clean, human-readable text in a consistent format.
This… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/LightOnOCR-bbox-mix-0126.lightonocr-pubtableslightonocr-pubtables-table-onlyLightOnOCR-bbox-bench
LightOnOCR-bbox-bench
Evaluation benchmark for assessing the ability of vision-language models (VLMs) to localize images within documents using bounding boxes. This dataset was introduced in the paper LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR.
Task Description
Given a document page (PDF), the model must predict bounding boxes around images (figures, charts, photographs, etc.) present in the document. This evaluates the model's… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/LightOnOCR-bbox-bench.samaritan_hebrew_LightOnOcr
Samaritan Hebrew OCR Dataset
Dataset Summary
The Samaritan Hebrew OCR Dataset is a specialized dataset for fine-tuning OCR models on Samaritan Hebrew manuscripts. This dataset contains 46,860 annotated samples extracted from 1,374 manuscript pages, converted from PAGE-XML format to the LightOnOCR-2 training format.
The dataset includes three types of samples:
Line-level samples: Individual textlines cropped using precise polygon masks (40,219 samples)
Paragraph-level… See the full description on the dataset page: https://huggingface.co/datasets/samaritan-ai/samaritan_hebrew_LightOnOcr.handbooks-lighton-ocr-32k-test
Document OCR using LightOnOCR-0.9B-32k-1025
This dataset contains OCR results from images in NationalLibraryOfScotland/Britain-and-UK-Handbooks-Dataset using LightOnOCR, a fast and compact 1B OCR model.
Processing Details
Source Dataset: NationalLibraryOfScotland/Britain-and-UK-Handbooks-Dataset
Model: lightonai/LightOnOCR-0.9B-32k-1025
Vocabulary Size: 32k tokens
Number of Samples: 4,096
Processing Time: 27.8 min
Processing Date: 2025-10-24 12:35 UTC… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/handbooks-lighton-ocr-32k-test.encyclopaedia_britannica_illustrated-lighton-ocr-32k-test
Document OCR using LightOnOCR-0.9B-32k-1025
This dataset contains OCR results from images in NationalLibraryOfScotland/Britain-and-UK-Handbooks-Dataset using LightOnOCR, a fast and compact 1B OCR model.
Processing Details
Source Dataset: NationalLibraryOfScotland/Britain-and-UK-Handbooks-Dataset
Model: lightonai/LightOnOCR-0.9B-32k-1025
Vocabulary Size: 32k tokens
Number of Samples: 100
Processing Time: 3.7 min
Processing Date: 2025-10-23 17:51 UTC… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/encyclopaedia_britannica_illustrated-lighton-ocr-32k-test.handbooks-lighton-ocr-32k-test-4
Document OCR using LightOnOCR-1B-1025
This dataset contains OCR results from images in davanstrien/handbooks-lighton-ocr-32k-test using LightOnOCR, a fast and compact 1B OCR model.
Processing Details
Source Dataset: davanstrien/handbooks-lighton-ocr-32k-test
Model: lightonai/LightOnOCR-1B-1025
Vocabulary Size: 151k tokens
Number of Samples: 512
Processing Time: 9.2 min
Processing Date: 2025-10-27 12:11 UTC
Configuration
Image Column: image
Output Column:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/handbooks-lighton-ocr-32k-test-4.lightonocr2-blind-spots
LightOnOCR-2-1B Blind Spots Dataset
Model Tested
lightonai/LightOnOCR-2-1B — a 1B parameter
end-to-end vision-language model for document OCR, released in 2026 under Apache 2.0.
How the Model Was Loaded
The model was loaded in Google Colab (free T4 GPU) using transformers installed from source:
import torch
from transformers import LightOnOcrForConditionalGeneration, LightOnOcrProcessor
device = "cuda"
dtype = torch.bfloat16
model =… See the full description on the dataset page: https://huggingface.co/datasets/Nonso-Analytics/lightonocr2-blind-spots.lighton-ocr2-test-v4
Document OCR using LightOnOCR-2-1B
This dataset contains OCR results from images in davanstrien/ufo-ColPali using LightOnOCR-2, a fast and compact 1B OCR model trained with RLVR.
Processing Details
Source Dataset: davanstrien/ufo-ColPali
Model: lightonai/LightOnOCR-2-1B
Number of Samples: 10
Processing Time: 2.8 min
Processing Date: 2026-01-29 17:49 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 16
Target… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/lighton-ocr2-test-v4.lighton-ocr2-test-v1
Document OCR using LightOnOCR-2-1B
This dataset contains OCR results from images in davanstrien/ufo-ColPali using LightOnOCR-2, a fast and compact 1B OCR model trained with RLVR.
Processing Details
Source Dataset: davanstrien/ufo-ColPali
Model: lightonai/LightOnOCR-2-1B
Number of Samples: 10
Processing Time: 2.9 min
Processing Date: 2026-01-29 14:30 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 16
Target… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/lighton-ocr2-test-v1.Misraj-DocOCR__run_LightOnOCR-2-1Bnls-highland-news-lighton-ocr2
Document OCR using LightOnOCR-2-1B
This dataset contains OCR results from images in davanstrien/nls-highland-news-sample using LightOnOCR-2, a fast and compact 1B OCR model trained with RLVR.
Processing Details
Source Dataset: davanstrien/nls-highland-news-sample
Model: lightonai/LightOnOCR-2-1B
Number of Samples: 68
Processing Time: 6.4 min
Processing Date: 2026-02-22 16:09 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/nls-highland-news-lighton-ocr2.fixtest-lighton-ocr2
Document OCR using LightOnOCR-2-1B
This dataset contains OCR results from images in davanstrien/ufo-ColPali using LightOnOCR-2, a fast and compact 1B OCR model trained with RLVR.
Processing Details
Source Dataset: davanstrien/ufo-ColPali
Model: lightonai/LightOnOCR-2-1B
Number of Samples: 2
Processing Time: 2.1 min
Processing Date: 2026-06-05 10:14 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 16… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/fixtest-lighton-ocr2.lightonocr-blindspots
LightOnOCR Blind Spots Datasets
Dataset Summary
This dataset contains examples where the OCR model
LightOnOCR-2-1B-base produces incorrect predictions.
The dataset was created to analyze failure cases and identify blind spots in the model.
Each dataset entry contains:
image_url - URL of the image
expected_output - Correct transcription of the image
model_output - Text generated by the model
The goal of this dataset is to help diagnose model weaknesses and guide future… See the full description on the dataset page: https://huggingface.co/datasets/usmanadaudu/lightonocr-blindspots.lightonocr-blindspots
