CoolFace
Datasetpublic

PiotrSty/wolne-lektury-polish-literature-corpus

Wolne Lektury Polish Literature Corpus Dataset Description A comprehensive corpus of Polish literary works from Wolne Lektury — a free digital library of public domain literature. All texts are in the public domain. The corpus was collected via the official REST API (https://wolnelektury.pl/api/), including full text of each work, metadata (author, epoch, genre, kind), and language information. Statistics Metric Value Records 7,316… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/wolne-lektury-polish-literature-corpus.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes116downloads
Dataset Card

Wolne Lektury Polish Literature Corpus

Dataset Description

A comprehensive corpus of Polish literary works from Wolne Lektury — a free digital library of public domain literature. All texts are in the public domain.

The corpus was collected via the official REST API (https://wolnelektury.pl/api/), including full text of each work, metadata (author, epoch, genre, kind), and language information.

Statistics

MetricValue
Records7,316
Authors572
Total characters319,880,903
Total words48,164,947
Estimated tokens (~1.3x words)62,614,431
LanguagesPolish (7,094), Lithuanian (79), German (54), English (47), French (27), Ukrainian (12), Hebrew (3)

Literary Coverage

  • —Epochs: Współczesność (2,507), Modernizm (1,190), Romantyzm (1,026), Dwudziestolecie międzywojenne (943), Pozytywizm (792), Renesans (495), Oświecenie (295), Barok (163), Starożytność (134), Średniowiecze (14)
  • —Kinds: Liryka (5,043), Epika (1,939), Dramat (336)
  • —Top genres: Wiersz (4,121), Bajka (417), Opowiadanie (393), Fraszka (316), Nowela (228), Powieść (198), Przypowieść (197), Pieśń (177), Dramat współczesny (173), Baśń (98)
  • —Top authors: Jan Kochanowski (453), Bolesław Leśmian (266), Maria Konopnicka (254), Ignacy Krasicki (223), K.I. Gałczyński (195), K.K. Baczyński (194), Stanisław Jachowicz (158), Adam Mickiewicz (149)

Fields

FieldTypeDescription
idstringSlug identifier
textstringFull text of the work
titlestringTitle
authorstringAuthor(s), comma-separated
translatorslist[string]Translator(s)
languagestringLanguage code (pol, lit, ger, eng, fre, ukr, heb)
epochslist[string]Literary epoch(s)
genreslist[string]Genre(s)
kindslist[string]Literary kind (Epika/Liryka/Dramat)
char_countintCharacter count
word_countintWord count
urlstringURL on wolnelektury.pl
has_audioboolHas audiobook available

Processing Pipeline

  1. 1.Scraping: Fetched book list from GET /api/books/ (7,630 books listed)
  2. 2.Metadata + text: For each book, fetched GET /api/books/{slug}/ for metadata and downloaded full text from txt URL
  3. 3.Deduplication: By ID and by text content (SHA256 hash of normalized text)
  4. 4.PII redaction: Phone numbers, emails, and PESEL numbers replaced with [PHONE], [EMAIL], [PESEL] placeholders (1,435 records affected)
  5. 5.Short text filtering: Texts with <10 words removed (0 removed — all texts are substantial)
  6. 6.Decontamination: N-gram overlap check against LLMzSzŁ, PoQuAD, and PES benchmark corpora — 0 contaminated records

Quality Assessment

  • —PII scan (post-redaction): Clean — 0 emails, 0 phones, 0 PESEL
  • —Decontamination: Clean — 0 contaminated records
  • —Verdict: Promising, ready for further evaluation

License

All works are in the Public Domain. The dataset compilation is also freely usable.

Source