datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
polish-dynaword
Polish DynaWord
A continuously developed, openly-licensed, human-text Polish corpus — a Polish
edition in the Dynaword
family (Enevoldsen et al., arXiv:2508.02271).
v0.2.5 stable · 4,319,200 documents · 9.64B tokens
(tiktoken proxy; canonical Llama-3 count at release) · 18 sources
Updated: 2026-08-14
v0.3-dev experimental track · quality/diversity workflow, source-gate
validation and candidate-data audits. This is development work, not a released
corpus version, and it does… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword.gollem-corpus-16b-pl
GoLLeM Corpus v4 PL — SlayerLab/gollem-corpus-16b-pl
The exact pretraining corpus of the Polish base model GoLLeM v4 (250M, trained from
scratch) — released before training, so the published bytes are byte-identical
(sha-tied) to what the model will see. Successor of
SlayerLab/gollem-corpus-2b-pl
(the v2/v3 corpus), scaled ~7.5x with per-record provenance this time.
16.58B unique tokens (GoLLeM V32k tokenizer, measured) =
~1.96B curated + 14.62B cleaned Polish web.
Per-record… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-16b-pl.gollem-corpus-2b-pl
GoLLeM Corpus 2B PL
Dokładny korpus treningowy polskiego modelu bazowego
SlayerLab/GoLLeM-110M-PL-v3
(oraz v2) — ten sam zbiór, po którym model przeszedł dwie epoki. Publikujemy go,
aby każdy mógł odtworzyć trening od zera na własnym tokenizerze.
Jak powstał ten plik. Korpus odzyskano przez zdekodowanie stokenizowanego
checkpointu treningowego (byte-level BPE dynaword-32k, round-trip bezstratny; granice
dokumentów = token <|endoftext|>). To jest dokładnie tekst, który model… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-2b-pl.hplt-v3-pl-cleaned
HPLT v3 Polish — Cleaned & PII-Gated
Oczyszczony, polski podzbiór korpusu web HPLT v3, przygotowany jako materiał pretreningowy. Autor / kurator zbioru: Arkadiusz Słota (SlayerLab). Wartość dodana względem surowego HPLT: wieloetapowa bramka PII (usuwanie numerów telefonów, identyfikatorów, adresów) z niezależną weryfikacją na pełnych danych + lekkie czyszczenie boilerplate.
Wersja: v1.0 — floor (bins 8_5 + 8_6). Track B (bins 9_1 + 8_1, po deduplikacji względem bazy dynaword)… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/hplt-v3-pl-cleaned.polish-dynaword-mix-extended-500M
Polish DynaWord Mix — Extended (~500M tokens)
Maintained by Arkadiusz Słota · SlayerLab
A curated, openly-licensed Polish text corpus for language-model pretraining and
research baselines. Built on top of the open polish-dynaword lineage and extended to
~500 million tokens (32k BPE) with additional curated, license-compatible sources
and a documented cleaning + PII-scrubbing pipeline.
Why this exists: most large Polish web corpora are legal/parliamentary-heavy and
carry… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword-mix-extended-500M.
