Lukaszl/pl-newspaper-pages-ocr-dataset-100
Polish Press Pages OCR Dataset 100 A small public image-only OCR dataset containing 100 Polish-language press / magazine-style page images sampled from a multilingual document collection. This dataset is designed as a lightweight evaluation sample for: OCR models VLM-based document understanding testing OCR robustness on Polish multi-column and press-style layouts Contents 100 JPG images Polish-language page images press / magazine-style layouts, including:… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-newspaper-pages-ocr-dataset-100.
Polish Press Pages OCR Dataset 100
A small public image-only OCR dataset containing 100 Polish-language press / magazine-style page images sampled from a multilingual document collection.
This dataset is designed as a lightweight evaluation sample for:
- OCR models
- VLM-based document understanding
- testing OCR robustness on Polish multi-column and press-style layouts
Contents
- 100 JPG images
- Polish-language page images
- press / magazine-style layouts, including:
- article pages
- multi-column text layouts
- pages with embedded images
- pages with side blocks and highlighted sections
- advertisement-style pages
- editorial / magazine-like compositions
What is NOT included
- ground-truth transcriptions
- OCR outputs
- benchmark scores
- dataset splits
Purpose
This is not a full benchmark.
It is a compact dataset sample intended for:
- testing OCR robustness on Polish press-style pages
- quick OCR sanity checks
- evaluating performance on non-trivial reading-order layouts
- checking how models handle mixed typography and page composition
Source
This dataset is a derived Polish-only subset of the original upstream dataset:
- AlekseyScorpi/docsonseverallanguages https://huggingface.co/datasets/AlekseyScorpi/docsonseverallanguages
The upstream dataset is multilingual. This repository contains a Polish-only image subset selected for OCR evaluation purposes.
License / attribution
This repository is a derived public subset of the upstream source.
Please refer to the original dataset page for licensing and attribution terms:
- https://huggingface.co/datasets/AlekseyScorpi/docsonseveral_languages
If you use this subset, please also attribute the original upstream dataset where appropriate.
Maintainer
clearOCR https://clearocr.com
Notes
- image-only dataset
- Polish-language subset
- press / magazine-style pages
- intended for evaluation, not training
- useful as a compact OCR stress-test sample
