CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lightonai /LightOnOCR-mix-0126 LightOnOCR-mix-0126 LightOnOCR-mix-0126 is a large-scale OCR training dataset built via distillation: a strong vision–language model is prompted to produce naturally ordered full-page transcriptions (Markdown with LaTeX math spans and HTML tables) from rendered document pages. The dataset is designed as supervision for end-to-end OCR / document-understanding models that aim to output clean, human-readable text in a consistent format. This repository releases the PDFA-derived… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/LightOnOCR-mix-0126.textimage-to-text10M<n<100M112 likes1k downloads8mo agoHugging Face02lightonai /LightOnOCR-bbox-mix-0126 LightOnOCR-bbox-mix-0126 LightOnOCR-bbox-mix-0126 is a large-scale OCR training dataset including layout information built via distillation: a strong vision–language model is prompted to produce naturally ordered full-page transcriptions (Markdown with LaTeX math spans and HTML tables) from rendered document pages. The dataset is designed as supervision for end-to-end OCR / document-understanding models that aim to output clean, human-readable text in a consistent format. This… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/LightOnOCR-bbox-mix-0126.texttext-to-image100K<n<1M12 likes261 downloads8mo agoHugging Face03mlflowai /lightonocr-pubtablesimage100K<n<1M0 likes251 downloads2mo agoHugging Face04mlflowai /lightonocr-pubtables-table-onlyimage100K<n<1M0 likes88 downloads2mo agoHugging Face05lightonai /LightOnOCR-bbox-bench LightOnOCR-bbox-bench Evaluation benchmark for assessing the ability of vision-language models (VLMs) to localize images within documents using bounding boxes. This dataset was introduced in the paper LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR. Task Description Given a document page (PDF), the model must predict bounding boxes around images (figures, charts, photographs, etc.) present in the document. This evaluates the model's… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/LightOnOCR-bbox-bench.object-detectionn<1K8 likes39 downloads8mo agoHugging Face06samaritan-ai /samaritan_hebrew_LightOnOcr Samaritan Hebrew OCR Dataset Dataset Summary The Samaritan Hebrew OCR Dataset is a specialized dataset for fine-tuning OCR models on Samaritan Hebrew manuscripts. This dataset contains 46,860 annotated samples extracted from 1,374 manuscript pages, converted from PAGE-XML format to the LightOnOCR-2 training format. The dataset includes three types of samples: Line-level samples: Individual textlines cropped using precise polygon masks (40,219 samples) Paragraph-level… See the full description on the dataset page: https://huggingface.co/datasets/samaritan-ai/samaritan_hebrew_LightOnOcr.image-to-text10K<n<100K0 likes22 downloads8mo agoHugging Face07davanstrien /handbooks-lighton-ocr-32k-test Document OCR using LightOnOCR-0.9B-32k-1025 This dataset contains OCR results from images in NationalLibraryOfScotland/Britain-and-UK-Handbooks-Dataset using LightOnOCR, a fast and compact 1B OCR model. Processing Details Source Dataset: NationalLibraryOfScotland/Britain-and-UK-Handbooks-Dataset Model: lightonai/LightOnOCR-0.9B-32k-1025 Vocabulary Size: 32k tokens Number of Samples: 4,096 Processing Time: 27.8 min Processing Date: 2025-10-24 12:35 UTC… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/handbooks-lighton-ocr-32k-test.image1K<n<10K0 likes21 downloads11mo agoHugging Face08davanstrien /encyclopaedia_britannica_illustrated-lighton-ocr-32k-test Document OCR using LightOnOCR-0.9B-32k-1025 This dataset contains OCR results from images in NationalLibraryOfScotland/Britain-and-UK-Handbooks-Dataset using LightOnOCR, a fast and compact 1B OCR model. Processing Details Source Dataset: NationalLibraryOfScotland/Britain-and-UK-Handbooks-Dataset Model: lightonai/LightOnOCR-0.9B-32k-1025 Vocabulary Size: 32k tokens Number of Samples: 100 Processing Time: 3.7 min Processing Date: 2025-10-23 17:51 UTC… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/encyclopaedia_britannica_illustrated-lighton-ocr-32k-test.imagen<1K2 likes17 downloads11mo agoHugging Face09davanstrien /handbooks-lighton-ocr-32k-test-4 Document OCR using LightOnOCR-1B-1025 This dataset contains OCR results from images in davanstrien/handbooks-lighton-ocr-32k-test using LightOnOCR, a fast and compact 1B OCR model. Processing Details Source Dataset: davanstrien/handbooks-lighton-ocr-32k-test Model: lightonai/LightOnOCR-1B-1025 Vocabulary Size: 151k tokens Number of Samples: 512 Processing Time: 9.2 min Processing Date: 2025-10-27 12:11 UTC Configuration Image Column: image Output Column:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/handbooks-lighton-ocr-32k-test-4.imagen<1K0 likes13 downloads11mo agoHugging Face10Nonso-Analytics /lightonocr2-blind-spots LightOnOCR-2-1B Blind Spots Dataset Model Tested lightonai/LightOnOCR-2-1B — a 1B parameter end-to-end vision-language model for document OCR, released in 2026 under Apache 2.0. How the Model Was Loaded The model was loaded in Google Colab (free T4 GPU) using transformers installed from source: import torch from transformers import LightOnOcrForConditionalGeneration, LightOnOcrProcessor device = "cuda" dtype = torch.bfloat16 model =… See the full description on the dataset page: https://huggingface.co/datasets/Nonso-Analytics/lightonocr2-blind-spots.textimage-to-textn<1K0 likes10 downloads7mo agoHugging Face11davanstrien /lighton-ocr2-test-v4 Document OCR using LightOnOCR-2-1B This dataset contains OCR results from images in davanstrien/ufo-ColPali using LightOnOCR-2, a fast and compact 1B OCR model trained with RLVR. Processing Details Source Dataset: davanstrien/ufo-ColPali Model: lightonai/LightOnOCR-2-1B Number of Samples: 10 Processing Time: 2.8 min Processing Date: 2026-01-29 17:49 UTC Configuration Image Column: image Output Column: markdown Dataset Split: train Batch Size: 16 Target… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/lighton-ocr2-test-v4.imagen<1K0 likes9 downloads8mo agoHugging Face12davanstrien /lighton-ocr2-test-v1 Document OCR using LightOnOCR-2-1B This dataset contains OCR results from images in davanstrien/ufo-ColPali using LightOnOCR-2, a fast and compact 1B OCR model trained with RLVR. Processing Details Source Dataset: davanstrien/ufo-ColPali Model: lightonai/LightOnOCR-2-1B Number of Samples: 10 Processing Time: 2.9 min Processing Date: 2026-01-29 14:30 UTC Configuration Image Column: image Output Column: markdown Dataset Split: train Batch Size: 16 Target… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/lighton-ocr2-test-v1.imagen<1K0 likes8 downloads8mo agoHugging Face13muhammad0-0hreden /Misraj-DocOCR__run_LightOnOCR-2-1Bimagen<1K0 likes7 downloads3mo agoHugging Face14davanstrien /nls-highland-news-lighton-ocr2 Document OCR using LightOnOCR-2-1B This dataset contains OCR results from images in davanstrien/nls-highland-news-sample using LightOnOCR-2, a fast and compact 1B OCR model trained with RLVR. Processing Details Source Dataset: davanstrien/nls-highland-news-sample Model: lightonai/LightOnOCR-2-1B Number of Samples: 68 Processing Time: 6.4 min Processing Date: 2026-02-22 16:09 UTC Configuration Image Column: image Output Column: markdown Dataset Split:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/nls-highland-news-lighton-ocr2.imagen<1K0 likes5 downloads7mo agoHugging Face15davanstrien /fixtest-lighton-ocr2 Document OCR using LightOnOCR-2-1B This dataset contains OCR results from images in davanstrien/ufo-ColPali using LightOnOCR-2, a fast and compact 1B OCR model trained with RLVR. Processing Details Source Dataset: davanstrien/ufo-ColPali Model: lightonai/LightOnOCR-2-1B Number of Samples: 2 Processing Time: 2.1 min Processing Date: 2026-06-05 10:14 UTC Configuration Image Column: image Output Column: markdown Dataset Split: train Batch Size: 16… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/fixtest-lighton-ocr2.imagen<1K0 likes5 downloads4mo agoHugging Face16usmanadaudu /lightonocr-blindspots LightOnOCR Blind Spots Datasets Dataset Summary This dataset contains examples where the OCR model LightOnOCR-2-1B-base produces incorrect predictions. The dataset was created to analyze failure cases and identify blind spots in the model. Each dataset entry contains: image_url - URL of the image expected_output - Correct transcription of the image model_output - Text generated by the model The goal of this dataset is to help diagnose model weaknesses and guide future… See the full description on the dataset page: https://huggingface.co/datasets/usmanadaudu/lightonocr-blindspots.textn<1K0 likes2 downloads6mo agoHugging Face17fiture99 /lightonocr-blindspots0 likes1 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.