CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SlayerLab /polish-dynaword Polish DynaWord A continuously developed, openly-licensed, human-text Polish corpus — a Polish edition in the Dynaword family (Enevoldsen et al., arXiv:2508.02271). v0.2.5 stable · 4,319,200 documents · 9.64B tokens (tiktoken proxy; canonical Llama-3 count at release) · 18 sources Updated: 2026-08-14 v0.3-dev experimental track · quality/diversity workflow, source-gate validation and candidate-data audits. This is development work, not a released corpus version, and it does… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword.text-generation1M<n<10M24 likes4.5k downloads3d agoHugging Face02SlayerLab /gollem-corpus-16b-pl GoLLeM Corpus v4 PL — SlayerLab/gollem-corpus-16b-pl The exact pretraining corpus of the Polish base model GoLLeM v4 (250M, trained from scratch) — released before training, so the published bytes are byte-identical (sha-tied) to what the model will see. Successor of SlayerLab/gollem-corpus-2b-pl (the v2/v3 corpus), scaled ~7.5x with per-record provenance this time. 16.58B unique tokens (GoLLeM V32k tokenizer, measured) = ~1.96B curated + 14.62B cleaned Polish web. Per-record… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-16b-pl.texttext-generation1M<n<10M1 likes344 downloads12d agoHugging Face03SlayerLab /gollem-corpus-2b-pl GoLLeM Corpus 2B PL Dokładny korpus treningowy polskiego modelu bazowego SlayerLab/GoLLeM-110M-PL-v3 (oraz v2) — ten sam zbiór, po którym model przeszedł dwie epoki. Publikujemy go, aby każdy mógł odtworzyć trening od zera na własnym tokenizerze. Jak powstał ten plik. Korpus odzyskano przez zdekodowanie stokenizowanego checkpointu treningowego (byte-level BPE dynaword-32k, round-trip bezstratny; granice dokumentów = token <|endoftext|>). To jest dokładnie tekst, który model… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-2b-pl.tabulartext-generation1M<n<10M1 likes294 downloads26d agoHugging Face04SlayerLab /hplt-v3-pl-cleaned HPLT v3 Polish — Cleaned & PII-Gated Oczyszczony, polski podzbiór korpusu web HPLT v3, przygotowany jako materiał pretreningowy. Autor / kurator zbioru: Arkadiusz Słota (SlayerLab). Wartość dodana względem surowego HPLT: wieloetapowa bramka PII (usuwanie numerów telefonów, identyfikatorów, adresów) z niezależną weryfikacją na pełnych danych + lekkie czyszczenie boilerplate. Wersja: v1.0 — floor (bins 8_5 + 8_6). Track B (bins 9_1 + 8_1, po deduplikacji względem bazy dynaword)… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/hplt-v3-pl-cleaned.texttext-generation10M<n<100M0 likes256 downloads1mo agoHugging Face05SlayerLab /polish-dynaword-mix-extended-500M Polish DynaWord Mix — Extended (~500M tokens) Maintained by Arkadiusz Słota · SlayerLab A curated, openly-licensed Polish text corpus for language-model pretraining and research baselines. Built on top of the open polish-dynaword lineage and extended to ~500 million tokens (32k BPE) with additional curated, license-compatible sources and a documented cleaning + PII-scrubbing pipeline. Why this exists: most large Polish web corpora are legal/parliamentary-heavy and carry… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword-mix-extended-500M.texttext-generation10M<n<100M0 likes99 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.