CoolFace
20 results

olmo

allenai /olmOCR-bench olmOCR-bench olmOCR-bench is a dataset of 1,403 PDF files, plus 7,010 unit test cases that capture properties of the output that a good OCR system should have. This benchmark evaluates the ability of OCR systems to accurately convert PDF documents to markdown format while preserving critical textual and structural information. Quick links: 📃 Paper 🛠️ Code 🎮 Demo Table 1. Distribution of Test Classes by Document Source Document Source Text Present Text… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-bench.document1K<n<10K290 likes48k downloads7mo agoHugging Faceallenai /olmo-mix-1124 OLMo 2 (November 2024) Pretraining set Collection of data used to train OLMo-2-1124 models. The majority of this dataset comes from DCLM-Baseline with no additional filtering, but we provide the explicit breakdowns below. Name Tokens Bytes (uncompressed) Documents License DCLM-Baseline 3.70T 21.3TB 2.95B CC-BY-4.0 Arxiv 20.8B 77.2GB 3.95M ODC-BY pes2o 58.6B 412GB 38M ODC-BY starcoder 83.0B 458GB 78.7M ODC-BY Algebraic-stack 11.8B 44.0GB 2.83M ODC-BY… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-mix-1124.texttext-generation1B<n<10B91 likes35k downloads1y agoHugging Faceshhdwi /olmocr-pre-rendered olmOCR-bench Pre-Rendered Pre-rendered PNG images of the olmOCR-bench benchmark dataset, ready for zero-setup evaluation of any OCR / vision model. What This Is The official olmOCR benchmark requires downloading 1,403 PDFs locally and rendering each page to a PNG image before sending it to a model. Every benchmark runner in the official repo does this same rendering step internally — see olmocr/data/renderpdf.py::render_pdf_to_base64png(). This dataset eliminates that… See the full description on the dataset page: https://huggingface.co/datasets/shhdwi/olmocr-pre-rendered.image1K<n<10K0 likes31k downloads7mo agoHugging Faceallenai /olmOCR-synthmix-1025 olmOCR-synthmix-1025 olmOCR-synthmix-1025 is a dataset of 2,186 single PDF pages, that have been synthetically rerendered into HTML by claude-sonnet-4-20250514. In total, across these PDF pages, 30,381 synthetic benchmark cases have been created, following the format of olmOCR-bench. These documents contain no overlap with the original olmOCR-bench documents, and thus can be used as RLVR training data to improve the performance of OCR engines. Directory Structure… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-synthmix-1025.document1K<n<10K3 likes10k downloads11mo agoHugging Faceredmoddata /dolmino_olmocr_pdfs0 likes4.5k downloads9mo agoHugging Faceallenai /OLMoE-mix-0924 OLMoE Mix (September 2024) The following data mix was used to train OLMoE-1B-7B, a Mixture-of-Experts LLM with 1B active and 7B total parameters released in September 2024. The base version of OLMoE-1B-7B can be found at this page, the SFT of OLMoE-1B-7B is available here, and a version combining SFT and DPO is available following this link. Statistics Subset Tokens Words Bytes Docs DCLM Baseline 1.0 3.86 T 3.38 T16.7 T 2.95 B Starcoder 101 B 63.9 B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/OLMoE-mix-0924.text-generation1B<n<10B57 likes4.3k downloads2y agoHugging Face