CoolFace
Datasetpublic

Lukaszl/pl-newspaper-pages-ocr-dataset-100

Polish Press Pages OCR Dataset 100 A small public image-only OCR dataset containing 100 Polish-language press / magazine-style page images sampled from a multilingual document collection. This dataset is designed as a lightweight evaluation sample for: OCR models VLM-based document understanding testing OCR robustness on Polish multi-column and press-style layouts Contents 100 JPG images Polish-language page images press / magazine-style layouts, including:… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-newspaper-pages-ocr-dataset-100.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes53downloads
Dataset Card

Polish Press Pages OCR Dataset 100

A small public image-only OCR dataset containing 100 Polish-language press / magazine-style page images sampled from a multilingual document collection.

This dataset is designed as a lightweight evaluation sample for:

  • —OCR models
  • —VLM-based document understanding
  • —testing OCR robustness on Polish multi-column and press-style layouts

Contents

  • —100 JPG images
  • —Polish-language page images
  • —press / magazine-style layouts, including:
  • —article pages
  • —multi-column text layouts
  • —pages with embedded images
  • —pages with side blocks and highlighted sections
  • —advertisement-style pages
  • —editorial / magazine-like compositions

What is NOT included

  • —ground-truth transcriptions
  • —OCR outputs
  • —benchmark scores
  • —dataset splits

Purpose

This is not a full benchmark.

It is a compact dataset sample intended for:

  • —testing OCR robustness on Polish press-style pages
  • —quick OCR sanity checks
  • —evaluating performance on non-trivial reading-order layouts
  • —checking how models handle mixed typography and page composition

Source

This dataset is a derived Polish-only subset of the original upstream dataset:

  • —AlekseyScorpi/docsonseverallanguages https://huggingface.co/datasets/AlekseyScorpi/docsonseverallanguages

The upstream dataset is multilingual. This repository contains a Polish-only image subset selected for OCR evaluation purposes.

License / attribution

This repository is a derived public subset of the upstream source.

Please refer to the original dataset page for licensing and attribution terms:

  • —https://huggingface.co/datasets/AlekseyScorpi/docsonseveral_languages

If you use this subset, please also attribute the original upstream dataset where appropriate.

Maintainer

clearOCR https://clearocr.com

Notes

  • —image-only dataset
  • —Polish-language subset
  • —press / magazine-style pages
  • —intended for evaluation, not training
  • —useful as a compact OCR stress-test sample