CoolFace
20 results

lightonocr

lightonai /LightOnOCR-mix-0126 LightOnOCR-mix-0126 LightOnOCR-mix-0126 is a large-scale OCR training dataset built via distillation: a strong vision–language model is prompted to produce naturally ordered full-page transcriptions (Markdown with LaTeX math spans and HTML tables) from rendered document pages. The dataset is designed as supervision for end-to-end OCR / document-understanding models that aim to output clean, human-readable text in a consistent format. This repository releases the PDFA-derived… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/LightOnOCR-mix-0126.textimage-to-text10M<n<100M112 likes851 downloads8mo agoHugging Facelightonai /LightOnOCR-bbox-mix-0126 LightOnOCR-bbox-mix-0126 LightOnOCR-bbox-mix-0126 is a large-scale OCR training dataset including layout information built via distillation: a strong vision–language model is prompted to produce naturally ordered full-page transcriptions (Markdown with LaTeX math spans and HTML tables) from rendered document pages. The dataset is designed as supervision for end-to-end OCR / document-understanding models that aim to output clean, human-readable text in a consistent format. This… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/LightOnOCR-bbox-mix-0126.texttext-to-image100K<n<1M12 likes243 downloads8mo agoHugging Facemlflowai /lightonocr-pubtablesimage100K<n<1M0 likes213 downloads2mo agoHugging Facemlflowai /lightonocr-pubtables-table-onlyimage100K<n<1M0 likes86 downloads2mo agoHugging Facesamaritan-ai /samaritan_hebrew_LightOnOcr Samaritan Hebrew OCR Dataset Dataset Summary The Samaritan Hebrew OCR Dataset is a specialized dataset for fine-tuning OCR models on Samaritan Hebrew manuscripts. This dataset contains 46,860 annotated samples extracted from 1,374 manuscript pages, converted from PAGE-XML format to the LightOnOCR-2 training format. The dataset includes three types of samples: Line-level samples: Individual textlines cropped using precise polygon masks (40,219 samples) Paragraph-level… See the full description on the dataset page: https://huggingface.co/datasets/samaritan-ai/samaritan_hebrew_LightOnOcr.image-to-text10K<n<100K0 likes42 downloads8mo agoHugging Facelightonai /LightOnOCR-bbox-bench LightOnOCR-bbox-bench Evaluation benchmark for assessing the ability of vision-language models (VLMs) to localize images within documents using bounding boxes. This dataset was introduced in the paper LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR. Task Description Given a document page (PDF), the model must predict bounding boxes around images (figures, charts, photographs, etc.) present in the document. This evaluates the model's… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/LightOnOCR-bbox-bench.object-detectionn<1K8 likes37 downloads8mo agoHugging Face