CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SlayerLab /gollem-corpus-16b-pl GoLLeM Corpus v4 PL — SlayerLab/gollem-corpus-16b-pl The exact pretraining corpus of the Polish base model GoLLeM v4 (250M, trained from scratch) — released before training, so the published bytes are byte-identical (sha-tied) to what the model will see. Successor of SlayerLab/gollem-corpus-2b-pl (the v2/v3 corpus), scaled ~7.5x with per-record provenance this time. 16.58B unique tokens (GoLLeM V32k tokenizer, measured) = ~1.96B curated + 14.62B cleaned Polish web. Per-record… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-16b-pl.texttext-generation1M<n<10M1 likes344 downloads12d agoHugging Face02SlayerLab /gollem-corpus-2b-pl GoLLeM Corpus 2B PL Dokładny korpus treningowy polskiego modelu bazowego SlayerLab/GoLLeM-110M-PL-v3 (oraz v2) — ten sam zbiór, po którym model przeszedł dwie epoki. Publikujemy go, aby każdy mógł odtworzyć trening od zera na własnym tokenizerze. Jak powstał ten plik. Korpus odzyskano przez zdekodowanie stokenizowanego checkpointu treningowego (byte-level BPE dynaword-32k, round-trip bezstratny; granice dokumentów = token <|endoftext|>). To jest dokładnie tekst, który model… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-2b-pl.tabulartext-generation1M<n<10M1 likes294 downloads26d agoHugging Face03SlayerLab /hplt-v3-pl-cleaned HPLT v3 Polish — Cleaned & PII-Gated Oczyszczony, polski podzbiór korpusu web HPLT v3, przygotowany jako materiał pretreningowy. Autor / kurator zbioru: Arkadiusz Słota (SlayerLab). Wartość dodana względem surowego HPLT: wieloetapowa bramka PII (usuwanie numerów telefonów, identyfikatorów, adresów) z niezależną weryfikacją na pełnych danych + lekkie czyszczenie boilerplate. Wersja: v1.0 — floor (bins 8_5 + 8_6). Track B (bins 9_1 + 8_1, po deduplikacji względem bazy dynaword)… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/hplt-v3-pl-cleaned.texttext-generation10M<n<100M0 likes256 downloads1mo agoHugging Face04SlayerLab /polish-dynaword-mix-extended-500M Polish DynaWord Mix — Extended (~500M tokens) Maintained by Arkadiusz Słota · SlayerLab A curated, openly-licensed Polish text corpus for language-model pretraining and research baselines. Built on top of the open polish-dynaword lineage and extended to ~500 million tokens (32k BPE) with additional curated, license-compatible sources and a documented cleaning + PII-scrubbing pipeline. Why this exists: most large Polish web corpora are legal/parliamentary-heavy and carry… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword-mix-extended-500M.texttext-generation10M<n<100M0 likes99 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.