slayer-lab
polish-dynaword
Polish DynaWord
A continuously developed, openly-licensed, human-text Polish corpus — a Polish
edition in the Dynaword
family (Enevoldsen et al., arXiv:2508.02271).
v0.2.5 stable · 4,319,200 documents · 9.64B tokens
(tiktoken proxy; canonical Llama-3 count at release) · 18 sources
Updated: 2026-08-14
v0.3-dev experimental track · quality/diversity workflow, source-gate
validation and candidate-data audits. This is development work, not a released
corpus version, and it does… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword.gollem-ui-mirrorminimal-en-corpus-5b
Minimal EN Corpus 5B
An English-language pretraining corpus prepared for controlled experiments with approximately 125M-parameter GPT-2 models based on karpathy/nanoGPT.
The name refers to the approximately 5B-token mixture selected before final BPE tokenization. With the included 12,288-token BPE tokenizer, the packaged nanoGPT training split contains 5,396,605,407 tokens.
Contents
The dataset provides both reusable source text and ready-to-train nanoGPT… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b.tokenizers
SlayerLab Tokenizers
Normalized tokenizer artifacts collected from the contributor directories in
slayerlabs/tokenizer,
pinned to source commit 1a5cd2c2e4df2287b4c19b3dbf5051f5d460fdc1.
The dataset contains one row per tokenizer: the 38 workshop submissions plus
the canonical SlayerLab Polish 32k tokenizer by kacperwikiel. Use the Dataset
Viewer to sort, filter, and compare tokenizers without navigating folders.
Columns
author: contributor's exact GitHub username… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/tokenizers.minimal-en-corpus-2.5b-v2
Minimal EN Corpus 2.5B 2.0
A deterministic English-language pretraining corpus for approximately 125M-parameter GPT-style models. This is the curated 2.0 generation of Minimal EN Corpus: source-pinned, model-free quality filtered, exact- and near-deduplicated, and decontaminated against the public evaluation splits of ARC, HellaSwag, PIQA, WinoGrande, OpenBookQA, BoolQ, MMLU, and GSM8K.
The name refers to the original approximately 2.5B-token source mixture. After cleaning… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-2.5b-v2.gollem-corpus-16b-pl
GoLLeM Corpus v4 PL — SlayerLab/gollem-corpus-16b-pl
The exact pretraining corpus of the Polish base model GoLLeM v4 (250M, trained from
scratch) — released before training, so the published bytes are byte-identical
(sha-tied) to what the model will see. Successor of
SlayerLab/gollem-corpus-2b-pl
(the v2/v3 corpus), scaled ~7.5x with per-record provenance this time.
16.58B unique tokens (GoLLeM V32k tokenizer, measured) =
~1.96B curated + 14.62B cleaned Polish web.
Per-record… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-16b-pl.
