PiotrSty/wolne-lektury-polish-literature-corpus
Wolne Lektury Polish Literature Corpus Dataset Description A comprehensive corpus of Polish literary works from Wolne Lektury — a free digital library of public domain literature. All texts are in the public domain. The corpus was collected via the official REST API (https://wolnelektury.pl/api/), including full text of each work, metadata (author, epoch, genre, kind), and language information. Statistics Metric Value Records 7,316… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/wolne-lektury-polish-literature-corpus.
Wolne Lektury Polish Literature Corpus
Dataset Description
A comprehensive corpus of Polish literary works from Wolne Lektury — a free digital library of public domain literature. All texts are in the public domain.
The corpus was collected via the official REST API (https://wolnelektury.pl/api/), including full text of each work, metadata (author, epoch, genre, kind), and language information.
Statistics
Literary Coverage
- Epochs: Współczesność (2,507), Modernizm (1,190), Romantyzm (1,026), Dwudziestolecie międzywojenne (943), Pozytywizm (792), Renesans (495), Oświecenie (295), Barok (163), Starożytność (134), Średniowiecze (14)
- Kinds: Liryka (5,043), Epika (1,939), Dramat (336)
- Top genres: Wiersz (4,121), Bajka (417), Opowiadanie (393), Fraszka (316), Nowela (228), Powieść (198), Przypowieść (197), Pieśń (177), Dramat współczesny (173), Baśń (98)
- Top authors: Jan Kochanowski (453), Bolesław Leśmian (266), Maria Konopnicka (254), Ignacy Krasicki (223), K.I. Gałczyński (195), K.K. Baczyński (194), Stanisław Jachowicz (158), Adam Mickiewicz (149)
Fields
Processing Pipeline
- Scraping: Fetched book list from
GET /api/books/(7,630 books listed) - Metadata + text: For each book, fetched
GET /api/books/{slug}/for metadata and downloaded full text fromtxtURL - Deduplication: By ID and by text content (SHA256 hash of normalized text)
- PII redaction: Phone numbers, emails, and PESEL numbers replaced with
[PHONE],[EMAIL],[PESEL]placeholders (1,435 records affected) - Short text filtering: Texts with <10 words removed (0 removed — all texts are substantial)
- Decontamination: N-gram overlap check against LLMzSzŁ, PoQuAD, and PES benchmark corpora — 0 contaminated records
Quality Assessment
- PII scan (post-redaction): Clean — 0 emails, 0 phones, 0 PESEL
- Decontamination: Clean — 0 contaminated records
- Verdict: Promising, ready for further evaluation
License
All works are in the Public Domain. The dataset compilation is also freely usable.
Source
- Website: https://wolnelektury.pl
- API: https://wolnelektury.pl/api/
