datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
polish-dynaword
Polish DynaWord
A continuously developed, openly-licensed, human-text Polish corpus — a Polish
edition in the Dynaword
family (Enevoldsen et al., arXiv:2508.02271).
v0.2.5 stable · 4,319,200 documents · 9.64B tokens
(tiktoken proxy; canonical Llama-3 count at release) · 18 sources
Updated: 2026-08-14
v0.3-dev experimental track · quality/diversity workflow, source-gate
validation and candidate-data audits. This is development work, not a released
corpus version, and it does… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword.gollem-ui-mirrorminimal-en-corpus-5b
Minimal EN Corpus 5B
An English-language pretraining corpus prepared for controlled experiments with approximately 125M-parameter GPT-2 models based on karpathy/nanoGPT.
The name refers to the approximately 5B-token mixture selected before final BPE tokenization. With the included 12,288-token BPE tokenizer, the packaged nanoGPT training split contains 5,396,605,407 tokens.
Contents
The dataset provides both reusable source text and ready-to-train nanoGPT… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b.gollem-corpus-16b-pl
GoLLeM Corpus v4 PL — SlayerLab/gollem-corpus-16b-pl
The exact pretraining corpus of the Polish base model GoLLeM v4 (250M, trained from
scratch) — released before training, so the published bytes are byte-identical
(sha-tied) to what the model will see. Successor of
SlayerLab/gollem-corpus-2b-pl
(the v2/v3 corpus), scaled ~7.5x with per-record provenance this time.
16.58B unique tokens (GoLLeM V32k tokenizer, measured) =
~1.96B curated + 14.62B cleaned Polish web.
Per-record… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-16b-pl.minimal-en-corpus-2.5b-v2
Minimal EN Corpus 2.5B 2.0
A deterministic English-language pretraining corpus for approximately 125M-parameter GPT-style models. This is the curated 2.0 generation of Minimal EN Corpus: source-pinned, model-free quality filtered, exact- and near-deduplicated, and decontaminated against the public evaluation splits of ARC, HellaSwag, PIQA, WinoGrande, OpenBookQA, BoolQ, MMLU, and GSM8K.
The name refers to the original approximately 2.5B-token source mixture. After cleaning… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-2.5b-v2.gollem-corpus-2b-pl
GoLLeM Corpus 2B PL
Dokładny korpus treningowy polskiego modelu bazowego
SlayerLab/GoLLeM-110M-PL-v3
(oraz v2) — ten sam zbiór, po którym model przeszedł dwie epoki. Publikujemy go,
aby każdy mógł odtworzyć trening od zera na własnym tokenizerze.
Jak powstał ten plik. Korpus odzyskano przez zdekodowanie stokenizowanego
checkpointu treningowego (byte-level BPE dynaword-32k, round-trip bezstratny; granice
dokumentów = token <|endoftext|>). To jest dokładnie tekst, który model… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-2b-pl.slayer-pl-8x3b
slayer-pl-8x3b
Osiem rozłącznych polskich zbiorów treningowych, po około 3 mld tokenów każdy (łącznie około 24 mld tokenów, licznik cl100k). Zbiory przeznaczone są do treningu wieloprzebiegowego: różnią się rozłącznym wycinkiem danych web (brak wspólnych dokumentów między zbiorami), natomiast współdzielą rdzeń zróżnicowany rejestrowo, warstwę prawno-urzędową oraz dosypkę angielską. Format wyjściowy: parquet. Tokenizacja pozostaje po stronie odbiorcy.
Kontrakt danych… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/slayer-pl-8x3b.hplt-v3-pl-cleaned
HPLT v3 Polish — Cleaned & PII-Gated
Oczyszczony, polski podzbiór korpusu web HPLT v3, przygotowany jako materiał pretreningowy. Autor / kurator zbioru: Arkadiusz Słota (SlayerLab). Wartość dodana względem surowego HPLT: wieloetapowa bramka PII (usuwanie numerów telefonów, identyfikatorów, adresów) z niezależną weryfikacją na pełnych danych + lekkie czyszczenie boilerplate.
Wersja: v1.0 — floor (bins 8_5 + 8_6). Track B (bins 9_1 + 8_1, po deduplikacji względem bazy dynaword)… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/hplt-v3-pl-cleaned.tokenizers
SlayerLab Tokenizers
Normalized tokenizer artifacts collected from the contributor directories in
slayerlabs/tokenizer,
pinned to source commit 1a5cd2c2e4df2287b4c19b3dbf5051f5d460fdc1.
The dataset contains one row per tokenizer: the 38 workshop submissions plus
the canonical SlayerLab Polish 32k tokenizer by kacperwikiel. Use the Dataset
Viewer to sort, filter, and compare tokenizers without navigating folders.
Columns
author: contributor's exact GitHub username… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/tokenizers.minimal-en-corpus-2.5b
Minimal EN Corpus 2.5B
An English-language pretraining corpus prepared for experiments with a roughly 125M-parameter GPT-2 model based on karpathy/nanoGPT.
The name refers to the approximately 2.5B-token mixture selected before final BPE tokenization. With the included 12,288-token BPE vocabulary, the packaged nanoGPT training split contains 2,689,323,439 tokens.
Contents
This dataset provides both reusable source text and ready-to-train nanoGPT binaries:… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-2.5b.fabryka-track-polish-mix
Fabryka Track Polish training mix (100 MB/source pack)
This dataset is the verified corpus pack used by track.fabryka.ai for training-pipeline tests.
It contains UTF-8 text samples plus one JSON metadata file per source and catalog.json.
The bounded sources were materialized from fixed Hugging Face revisions. bytes in the catalog is the exact UTF-8 byte size; the current Track byte-token trainer counts one byte as one training token.
This is a reproducible workflow pack, not a… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/fabryka-track-polish-mix.polish-dynaword-mix-extended-500M
Polish DynaWord Mix — Extended (~500M tokens)
Maintained by Arkadiusz Słota · SlayerLab
A curated, openly-licensed Polish text corpus for language-model pretraining and
research baselines. Built on top of the open polish-dynaword lineage and extended to
~500 million tokens (32k BPE) with additional curated, license-compatible sources
and a documented cleaning + PII-scrubbing pipeline.
Why this exists: most large Polish web corpora are legal/parliamentary-heavy and
carry… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword-mix-extended-500M.gollem-v5-arcmix-9b
GoLLeM-v5 ARC-MIX 9.4B (EN)
An English pretraining corpus for the GoLLeM-v5 tiny-LM efficiency track (Glint Tiny-ML Leaderboard).
ARC-MIX is a reasoning/knowledge-enriched expansion of
SlayerLab/minimal-en-corpus-5b:
the base multi-source EN mixture with ARC-relevant data upweighted (gold ~3×, related ~2×)
to strengthen the ARC-Easy axis — the binding efficiency-constraint at this scale.
Key facts
Property
Value
Tokens (packaged, uint16)
9,417,035,832… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-v5-arcmix-9b.lojban-master
Lojban Master Corpus
Rights review in progress. This corpus combines 15 sources and does not
yet contain a complete source-by-source map of licences, attribution duties,
redistribution permissions and personal-data considerations. license: other is a warning, not a blanket licence. Do not redistribute the combined
corpus or use it for commercial model training until the applicable rights for
every selected row/source have been verified. Contact: k.wikiel@gmail.com.
The largest… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/lojban-master.gollem-v5-expand
