olmo
Datasets
All datasets matching “olmo”olmOCR-bench
olmOCR-bench
olmOCR-bench is a dataset of 1,403 PDF files, plus 7,010 unit test cases that capture properties of the output that a good OCR system should have.
This benchmark evaluates the ability of OCR systems to accurately convert PDF documents to markdown format while preserving critical textual and structural information.
Quick links:
📃 Paper
🛠️ Code
🎮 Demo
Table 1. Distribution of Test Classes by Document Source
Document Source
Text Present
Text… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-bench.olmo-mix-1124
OLMo 2 (November 2024) Pretraining set
Collection of data used to train OLMo-2-1124 models. The majority of this dataset comes from DCLM-Baseline with no additional filtering, but we provide the explicit breakdowns below.
Name
Tokens
Bytes (uncompressed)
Documents
License
DCLM-Baseline
3.70T
21.3TB
2.95B
CC-BY-4.0
Arxiv
20.8B
77.2GB
3.95M
ODC-BY
pes2o
58.6B
412GB
38M
ODC-BY
starcoder
83.0B
458GB
78.7M
ODC-BY
Algebraic-stack
11.8B
44.0GB
2.83M
ODC-BY… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-mix-1124.olmocr-pre-rendered
olmOCR-bench Pre-Rendered
Pre-rendered PNG images of the olmOCR-bench benchmark dataset, ready for zero-setup evaluation of any OCR / vision model.
What This Is
The official olmOCR benchmark requires downloading 1,403 PDFs locally and rendering each page to a PNG image before sending it to a model. Every benchmark runner in the official repo does this same rendering step internally — see olmocr/data/renderpdf.py::render_pdf_to_base64png().
This dataset eliminates that… See the full description on the dataset page: https://huggingface.co/datasets/shhdwi/olmocr-pre-rendered.olmOCR-synthmix-1025
olmOCR-synthmix-1025
olmOCR-synthmix-1025 is a dataset of 2,186 single PDF pages, that have been synthetically rerendered into HTML by
claude-sonnet-4-20250514.
In total, across these PDF pages, 30,381 synthetic benchmark cases have been created, following the format of olmOCR-bench.
These documents contain no overlap with the original olmOCR-bench documents, and thus can be used as RLVR training
data to improve the performance of OCR engines.
Directory Structure… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-synthmix-1025.dolmino_olmocr_pdfsOLMoE-mix-0924
OLMoE Mix (September 2024)
The following data mix was used to train OLMoE-1B-7B, a Mixture-of-Experts LLM with 1B active and 7B total parameters released in September 2024.
The base version of OLMoE-1B-7B can be found at this page, the SFT of OLMoE-1B-7B is available here, and a version combining SFT and DPO is available following this link.
Statistics
Subset
Tokens
Words
Bytes
Docs
DCLM Baseline 1.0
3.86 T
3.38 T16.7 T
2.95 B
Starcoder
101 B
63.9 B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/OLMoE-mix-0924.
