datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
olmocr-pre-rendered
olmOCR-bench Pre-Rendered
Pre-rendered PNG images of the olmOCR-bench benchmark dataset, ready for zero-setup evaluation of any OCR / vision model.
What This Is
The official olmOCR benchmark requires downloading 1,403 PDFs locally and rendering each page to a PNG image before sending it to a model. Every benchmark runner in the official repo does this same rendering step internally — see olmocr/data/renderpdf.py::render_pdf_to_base64png().
This dataset eliminates that… See the full description on the dataset page: https://huggingface.co/datasets/shhdwi/olmocr-pre-rendered.olmoearth_pretrain_datasetThis is the pre-training dataset for training the OlmoEarth pre-trained remote sensing foundation models.
Documentation is on GitHub at https://github.com/allenai/olmoearth_pretrain/blob/main/docs/Pretraining-Dataset.md
The dataset is released under CC BY 4.0. It includes data from the following sources:
Sentinel-2 L2A imagery from the European Space Agency, available under the Copernicus Sentinel Data and Service Legal Notice
Sentinel-1 GRD IW vv+vh imagery from the European Space Agency… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmoearth_pretrain_dataset.olmOCR_bench
Dataset Card for olmocr-bench
This is a FiftyOne dataset with 7019 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/olmOCR_bench")
# Launch the App
session = fo.launch_app(dataset)
Here is the completed dataset card, filled in… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/olmOCR_bench.olmoearth_lcc
OlmoEarth Land Cover Change (LCC) Dataset
This dataset contains point-based annotations of land cover change, used to
train the OlmoEarth LCC model (https://olmoearth-lcc.allen.ai). The model
detects recent land cover change from Sentinel-2 time series: it inputs a Sentinel-2
image time series with 16 quarterly images (to establish a historical baseline) and 4
recent biweekly images (to detect changes soon after they occur) and predicts, per pixel,
whether a land cover change… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmoearth_lcc.olmOCR-mix-1025-Photoreal
⚠️ Important Feedback Invitation
If this dataset has brought stronger positive or negative improvements to your model training,
I warmly welcome you to email me your feedback at any time.
This is very important to me — thank you!
📧Contact: hi@support.alrowilde.com
Photorealistic enhancements of document pages from allenai/olmOCR-mix-1025.Clean PDF renders are transformed into realistic scanned/photographed images with natural lighting, paper texture, shadows and capture… See the full description on the dataset page: https://huggingface.co/datasets/AlroWilde/olmOCR-mix-1025-Photoreal.portuguese-eval-logs-olmo2-smollm3
Evaluation Logs on Portuguese Benchmarks for OLMo-2 and SmolLM3
These logs contain benchmark results across a suite of Portuguese-language tasks. The data consists of recordings of the performance of various 3 different models at different checkpoints throughout their pretraining runs:
SmolLM3
OLMo-2-0425-1B
OLMo-2-1124-7B
Splits
Each split (smollm3_3b, olmo2_1b, olmo2_7b) contains rows for model checkpoints and columns for benchmark scores (e.g., ASSIN2 RTE, ENEM… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-eval-logs-olmo2-smollm3.finevision-olmocr-processedolmoearth_projects_awfThis model is for fine-tuning OlmoEarth-v1-Base to map land use and land cover in southern Kenya.
The data is annotated by experts at the African Wildlife Foundation.
For more details, please see the documentation on GitHub at https://github.com/allenai/olmoearth_projects/blob/main/docs/awf.md
olmOCR-mix-0225-documents_cleaned
olmOCR-mix-0225-documents_cleaned
The olmOCR-mix-0225-documents__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
156,730
QA turns
569,306
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
127
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/olmOCR-mix-0225-documents_cleaned.olmo-ocrolmOCR-mix-0225-documents_cleaned
olmOCR-mix-0225-documents_cleaned
The olmOCR-mix-0225-documents__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
156,730
QA turns
569,306
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
127
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/olmOCR-mix-0225-documents_cleaned.olmOCR-mix-0225-books_cleaned
olmOCR-mix-0225-books_cleaned
The olmOCR-mix-0225-books__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
15,022
QA turns
46,049
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
13
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/olmOCR-mix-0225-books_cleaned.test_olmocr_documents_qatest-olmocr2
Document OCR using olmOCR-2-7B-1025-FP8
This dataset contains markdown-formatted OCR results from images in davanstrien/test-olmocr2 using olmOCR-2-7B.
Processing Details
Source Dataset: davanstrien/test-olmocr2
Model: allenai/olmOCR-2-7B-1025-FP8
Number of Samples: 100
Processing Time: 0h 3m 32s
Processing Date: 2025-10-23 17:00 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 512
Max Model Length: 16,384… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/test-olmocr2.olmocr2-issue14-test
Document OCR using olmOCR-2-7B-1025-FP8
This dataset contains markdown-formatted OCR results from images in davanstrien/ufo-ColPali using olmOCR-2-7B.
Processing Details
Source Dataset: davanstrien/ufo-ColPali
Model: allenai/olmOCR-2-7B-1025-FP8
Number of Samples: 5
Processing Time: 0h 3m 27s
Processing Date: 2026-06-05 12:45 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 16
Max Model Length: 16,384… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/olmocr2-issue14-test.marker_benchmark_comparison_olmocr_llmhumatheque-vlm-pred-olmocr2olmOCR-mix-0225-books_cleaned
olmOCR-mix-0225-books_cleaned
The olmOCR-mix-0225-books__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
15,022
QA turns
46,049
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
13
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/olmOCR-mix-0225-books_cleaned.ocr-thesis-surya-marker-qwens-olmOcrtesting_olmocr2
Dataset Card for Voxel51/consolidated_receipt_dataset
This is a FiftyOne dataset with 800 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/testing_olmocr2")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/testing_olmocr2.olm-ocr-datasetprocessed_olmocr_vllmnewspapers-olmocr2
Document OCR using olmOCR-2-7B-1025-FP8
This dataset contains markdown-formatted OCR results from images in davanstrien/newspapers-with-images using olmOCR-2-7B.
Processing Details
Source Dataset: davanstrien/newspapers-with-images
Model: allenai/olmOCR-2-7B-1025-FP8
Number of Samples: 100
Processing Time: 0h 9m 29s
Processing Date: 2025-10-22 18:57 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 4… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/newspapers-olmocr2.olmOCR-mix-0225-documents-smallolmocr220k-with-images
