datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PL-Guard
Dataset Summary
Released PL-Guard dataset consists of two splits, test and test_adversarial. Here's the breakdown:
test used to evaluate safety classifiers.
test_adversarial used to evaluate safety classifiers on perturbated samples derived from the test.
More detailed information is available in the publication.
Usage
from datasets import load_dataset
# Load the test dataset
dataset = load_dataset("NASK-PIB/PL-Guard",'test')
# Load the test_adversarial dataset… See the full description on the dataset page: https://huggingface.co/datasets/NASK-PIB/PL-Guard.pl-github-readmes
Polish GitHub README Corpus — Polskie README z repozytoriów GitHub
Korpus polskojęzycznych plików README z repozytoriów GitHub, wygenerowany z GitHub API na podstawie klasyfikacji językowych datasetu github/multilingual-repositories (CC0-1.0).
Statystyki
Metryka
Wartość
Repozytoria (klasyfikacja PL)
209,519
Pobrane README (90 min scrape)
2,313
Po deduplikacji i filtrowaniu
2,203
Znaki
54,412,484
Słowa
6,605,910
Tokeny (szac.)
~8,917,978… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/pl-github-readmes.pl-government-docs-mix-ocr-dataset-v1-results
OCR Bench Results: Polish government documents benchmark
VLM-as-judge pairwise evaluation of OCR models on a dataset of real Polish government and public administration documents.
This benchmark focuses on structured, text-heavy documents typical for public institutions, including official forms, templates, administrative documents, and scanned materials.
As with all OCR benchmarks, results are document-type specific and should not be interpreted as a universal ranking across all… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-government-docs-mix-ocr-dataset-v1-results.pl-gear-help Content from gear-help
