plg
Datasets
All datasets matching “plg”pl-government-docs-mix-ocr-dataset
Polish municipal administrative documents OCR dataset
An OCR-oriented image dataset built from publicly available Polish municipal administrative documents.
The dataset contains pages collected from materials such as:
resolutions (uchwały)
ordinances / regulations
annexes
official notices
tabular administrative pages
stamped and signed office documents
other municipal administrative materials
All files in this repository are provided as images only.
What is inside… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-government-docs-mix-ocr-dataset.pl-github-readmes
Polish GitHub README Corpus — Polskie README z repozytoriów GitHub
Korpus polskojęzycznych plików README z repozytoriów GitHub, wygenerowany z GitHub API na podstawie klasyfikacji językowych datasetu github/multilingual-repositories (CC0-1.0).
Statystyki
Metryka
Wartość
Repozytoria (klasyfikacja PL)
209,519
Pobrane README (90 min scrape)
2,313
Po deduplikacji i filtrowaniu
2,203
Znaki
54,412,484
Słowa
6,605,910
Tokeny (szac.)
~8,917,978… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/pl-github-readmes.PL-Guard
Dataset Summary
Released PL-Guard dataset consists of two splits, test and test_adversarial. Here's the breakdown:
test used to evaluate safety classifiers.
test_adversarial used to evaluate safety classifiers on perturbated samples derived from the test.
More detailed information is available in the publication.
Usage
from datasets import load_dataset
# Load the test dataset
dataset = load_dataset("NASK-PIB/PL-Guard",'test')
# Load the test_adversarial dataset… See the full description on the dataset page: https://huggingface.co/datasets/NASK-PIB/PL-Guard.pl-government-docs-mix-ocr-dataset-v1-results
OCR Bench Results: Polish government documents benchmark
VLM-as-judge pairwise evaluation of OCR models on a dataset of real Polish government and public administration documents.
This benchmark focuses on structured, text-heavy documents typical for public institutions, including official forms, templates, administrative documents, and scanned materials.
As with all OCR benchmarks, results are document-type specific and should not be interpreted as a universal ranking across all… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-government-docs-mix-ocr-dataset-v1-results.pl-government-docs-mix-ocr-dataset-v1
Document OCR using GLM-OCR
This dataset contains OCR results from images in Lukaszl/pl-government-docs-mix-ocr-dataset using GLM-OCR, a compact 0.9B OCR model achieving SOTA performance.
Processing Details
Source Dataset: Lukaszl/pl-government-docs-mix-ocr-dataset
Model: zai-org/GLM-OCR
Task: text recognition
Number of Samples: 169
Processing Time: 9.2 min
Processing Date: 2026-04-03 11:18 UTC
Configuration
Image Column: image
Output Column: markdown… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-government-docs-mix-ocr-dataset-v1.plg-playbook
PLG Playbook — Product-Led Growth Implementation
The complete Product-Led Growth playbook covering freemium model design, self-serve onboarding, activation metrics, and the transition from individual users...
📦 Install on ClawHub
clawhub install plg-playbook
Then ask your AI agent:
"Design a freemium tier that doesn't cannibalize our $99/mo paid plan"
Installs the full PLG Playbook — Product-Led Growth Implementation playbook — battle-tested with 30+ Product… See the full description on the dataset page: https://huggingface.co/datasets/Gingiris/plg-playbook.
