CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SlayerLab /polish-dynaword Polish DynaWord A continuously developed, openly-licensed, human-text Polish corpus — a Polish edition in the Dynaword family (Enevoldsen et al., arXiv:2508.02271). v0.2.5 stable · 4,319,200 documents · 9.64B tokens (tiktoken proxy; canonical Llama-3 count at release) · 18 sources Updated: 2026-08-14 v0.3-dev experimental track · quality/diversity workflow, source-gate validation and candidate-data audits. This is development work, not a released corpus version, and it does… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword.text-generation1M<n<10M24 likes4.5k downloads3d agoHugging Face02SlayerLab /gollem-ui-mirror0 likes3.6k downloads3m agoHugging Face03SlayerLab /minimal-en-corpus-5b Minimal EN Corpus 5B An English-language pretraining corpus prepared for controlled experiments with approximately 125M-parameter GPT-2 models based on karpathy/nanoGPT. The name refers to the approximately 5B-token mixture selected before final BPE tokenization. With the included 12,288-token BPE tokenizer, the packaged nanoGPT training split contains 5,396,605,407 tokens. Contents The dataset provides both reusable source text and ready-to-train nanoGPT… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b.text1M<n<10M1 likes2.4k downloads3d agoHugging Face04SlayerLab /gollem-corpus-16b-pl GoLLeM Corpus v4 PL — SlayerLab/gollem-corpus-16b-pl The exact pretraining corpus of the Polish base model GoLLeM v4 (250M, trained from scratch) — released before training, so the published bytes are byte-identical (sha-tied) to what the model will see. Successor of SlayerLab/gollem-corpus-2b-pl (the v2/v3 corpus), scaled ~7.5x with per-record provenance this time. 16.58B unique tokens (GoLLeM V32k tokenizer, measured) = ~1.96B curated + 14.62B cleaned Polish web. Per-record… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-16b-pl.texttext-generation1M<n<10M1 likes344 downloads12d agoHugging Face05SlayerLab /minimal-en-corpus-2.5b-v2 Minimal EN Corpus 2.5B 2.0 A deterministic English-language pretraining corpus for approximately 125M-parameter GPT-style models. This is the curated 2.0 generation of Minimal EN Corpus: source-pinned, model-free quality filtered, exact- and near-deduplicated, and decontaminated against the public evaluation splits of ARC, HellaSwag, PIQA, WinoGrande, OpenBookQA, BoolQ, MMLU, and GSM8K. The name refers to the original approximately 2.5B-token source mixture. After cleaning… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-2.5b-v2.text1M<n<10M0 likes342 downloads3d agoHugging Face06SlayerLab /gollem-corpus-2b-pl GoLLeM Corpus 2B PL Dokładny korpus treningowy polskiego modelu bazowego SlayerLab/GoLLeM-110M-PL-v3 (oraz v2) — ten sam zbiór, po którym model przeszedł dwie epoki. Publikujemy go, aby każdy mógł odtworzyć trening od zera na własnym tokenizerze. Jak powstał ten plik. Korpus odzyskano przez zdekodowanie stokenizowanego checkpointu treningowego (byte-level BPE dynaword-32k, round-trip bezstratny; granice dokumentów = token <|endoftext|>). To jest dokładnie tekst, który model… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-2b-pl.tabulartext-generation1M<n<10M1 likes294 downloads26d agoHugging Face07SlayerLab /slayer-pl-8x3b slayer-pl-8x3b Osiem rozłącznych polskich zbiorów treningowych, po około 3 mld tokenów każdy (łącznie około 24 mld tokenów, licznik cl100k). Zbiory przeznaczone są do treningu wieloprzebiegowego: różnią się rozłącznym wycinkiem danych web (brak wspólnych dokumentów między zbiorami), natomiast współdzielą rdzeń zróżnicowany rejestrowo, warstwę prawno-urzędową oraz dosypkę angielską. Format wyjściowy: parquet. Tokenizacja pozostaje po stronie odbiorcy. Kontrakt danych… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/slayer-pl-8x3b.text10M<n<100M0 likes278 downloads29d agoHugging Face08SlayerLab /hplt-v3-pl-cleaned HPLT v3 Polish — Cleaned & PII-Gated Oczyszczony, polski podzbiór korpusu web HPLT v3, przygotowany jako materiał pretreningowy. Autor / kurator zbioru: Arkadiusz Słota (SlayerLab). Wartość dodana względem surowego HPLT: wieloetapowa bramka PII (usuwanie numerów telefonów, identyfikatorów, adresów) z niezależną weryfikacją na pełnych danych + lekkie czyszczenie boilerplate. Wersja: v1.0 — floor (bins 8_5 + 8_6). Track B (bins 9_1 + 8_1, po deduplikacji względem bazy dynaword)… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/hplt-v3-pl-cleaned.texttext-generation10M<n<100M0 likes256 downloads1mo agoHugging Face09SlayerLab /tokenizers SlayerLab Tokenizers Normalized tokenizer artifacts collected from the contributor directories in slayerlabs/tokenizer, pinned to source commit 1a5cd2c2e4df2287b4c19b3dbf5051f5d460fdc1. The dataset contains one row per tokenizer: the 38 workshop submissions plus the canonical SlayerLab Polish 32k tokenizer by kacperwikiel. Use the Dataset Viewer to sort, filter, and compare tokenizers without navigating folders. Columns author: contributor's exact GitHub username… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/tokenizers.tabularn<1K0 likes249 downloads2d agoHugging Face10SlayerLab /minimal-en-corpus-2.5b Minimal EN Corpus 2.5B An English-language pretraining corpus prepared for experiments with a roughly 125M-parameter GPT-2 model based on karpathy/nanoGPT. The name refers to the approximately 2.5B-token mixture selected before final BPE tokenization. With the included 12,288-token BPE vocabulary, the packaged nanoGPT training split contains 2,689,323,439 tokens. Contents This dataset provides both reusable source text and ready-to-train nanoGPT binaries:… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-2.5b.text1M<n<10M0 likes193 downloads28d agoHugging Face11SlayerLab /fabryka-track-polish-mix Fabryka Track Polish training mix (100 MB/source pack) This dataset is the verified corpus pack used by track.fabryka.ai for training-pipeline tests. It contains UTF-8 text samples plus one JSON metadata file per source and catalog.json. The bounded sources were materialized from fixed Hugging Face revisions. bytes in the catalog is the exact UTF-8 byte size; the current Track byte-token trainer counts one byte as one training token. This is a reproducible workflow pack, not a… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/fabryka-track-polish-mix.tabularn<1K1 likes161 downloads17d agoHugging Face12SlayerLab /polish-dynaword-mix-extended-500M Polish DynaWord Mix — Extended (~500M tokens) Maintained by Arkadiusz Słota · SlayerLab A curated, openly-licensed Polish text corpus for language-model pretraining and research baselines. Built on top of the open polish-dynaword lineage and extended to ~500 million tokens (32k BPE) with additional curated, license-compatible sources and a documented cleaning + PII-scrubbing pipeline. Why this exists: most large Polish web corpora are legal/parliamentary-heavy and carry… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword-mix-extended-500M.texttext-generation10M<n<100M0 likes99 downloads1mo agoHugging Face13SlayerLab /gollem-v5-arcmix-9b GoLLeM-v5 ARC-MIX 9.4B (EN) An English pretraining corpus for the GoLLeM-v5 tiny-LM efficiency track (Glint Tiny-ML Leaderboard). ARC-MIX is a reasoning/knowledge-enriched expansion of SlayerLab/minimal-en-corpus-5b: the base multi-source EN mixture with ARC-relevant data upweighted (gold ~3×, related ~2×) to strengthen the ARC-Easy axis — the binding efficiency-constraint at this scale. Key facts Property Value Tokens (packaged, uint16) 9,417,035,832… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-v5-arcmix-9b.1M<n<10M1 likes34 downloads2d agoHugging Face14SlayerLab /lojban-master Lojban Master Corpus Rights review in progress. This corpus combines 15 sources and does not yet contain a complete source-by-source map of licences, attribution duties, redistribution permissions and personal-data considerations. license: other is a warning, not a blanket licence. Do not redistribute the combined corpus or use it for commercial model training until the applicable rights for every selected row/source have been verified. Contact: k.wikiel@gmail.com. The largest… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/lojban-master.text100K<n<1M0 likes32 downloads2mo agoHugging Face15SlayerLab /gollem-v5-expand0 likes31 downloads4d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.