datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gollem-corpus-2b-pl
GoLLeM Corpus 2B PL
Dokładny korpus treningowy polskiego modelu bazowego
SlayerLab/GoLLeM-110M-PL-v3
(oraz v2) — ten sam zbiór, po którym model przeszedł dwie epoki. Publikujemy go,
aby każdy mógł odtworzyć trening od zera na własnym tokenizerze.
Jak powstał ten plik. Korpus odzyskano przez zdekodowanie stokenizowanego
checkpointu treningowego (byte-level BPE dynaword-32k, round-trip bezstratny; granice
dokumentów = token <|endoftext|>). To jest dokładnie tekst, który model… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-2b-pl.tokenizers
SlayerLab Tokenizers
Normalized tokenizer artifacts collected from the contributor directories in
slayerlabs/tokenizer,
pinned to source commit 1a5cd2c2e4df2287b4c19b3dbf5051f5d460fdc1.
The dataset contains one row per tokenizer: the 38 workshop submissions plus
the canonical SlayerLab Polish 32k tokenizer by kacperwikiel. Use the Dataset
Viewer to sort, filter, and compare tokenizers without navigating folders.
Columns
author: contributor's exact GitHub username… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/tokenizers.fabryka-track-polish-mix
Fabryka Track Polish training mix (100 MB/source pack)
This dataset is the verified corpus pack used by track.fabryka.ai for training-pipeline tests.
It contains UTF-8 text samples plus one JSON metadata file per source and catalog.json.
The bounded sources were materialized from fixed Hugging Face revisions. bytes in the catalog is the exact UTF-8 byte size; the current Track byte-token trainer counts one byte as one training token.
This is a reproducible workflow pack, not a… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/fabryka-track-polish-mix.
