CoolFace
15 results

slayer-lab

SlayerLab /polish-dynaword Polish DynaWord A continuously developed, openly-licensed, human-text Polish corpus — a Polish edition in the Dynaword family (Enevoldsen et al., arXiv:2508.02271). v0.2.5 stable · 4,319,200 documents · 9.64B tokens (tiktoken proxy; canonical Llama-3 count at release) · 18 sources Updated: 2026-08-14 v0.3-dev experimental track · quality/diversity workflow, source-gate validation and candidate-data audits. This is development work, not a released corpus version, and it does… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword.text-generation1M<n<10M24 likes4.6k downloads3d agoHugging FaceSlayerLab /gollem-ui-mirror0 likes2.8k downloads7h agoHugging FaceSlayerLab /minimal-en-corpus-5b Minimal EN Corpus 5B An English-language pretraining corpus prepared for controlled experiments with approximately 125M-parameter GPT-2 models based on karpathy/nanoGPT. The name refers to the approximately 5B-token mixture selected before final BPE tokenization. With the included 12,288-token BPE tokenizer, the packaged nanoGPT training split contains 5,396,605,407 tokens. Contents The dataset provides both reusable source text and ready-to-train nanoGPT… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b.text1M<n<10M1 likes2.3k downloads2d agoHugging FaceSlayerLab /tokenizers SlayerLab Tokenizers Normalized tokenizer artifacts collected from the contributor directories in slayerlabs/tokenizer, pinned to source commit 1a5cd2c2e4df2287b4c19b3dbf5051f5d460fdc1. The dataset contains one row per tokenizer: the 38 workshop submissions plus the canonical SlayerLab Polish 32k tokenizer by kacperwikiel. Use the Dataset Viewer to sort, filter, and compare tokenizers without navigating folders. Columns author: contributor's exact GitHub username… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/tokenizers.tabularn<1K0 likes385 downloads2d agoHugging FaceSlayerLab /minimal-en-corpus-2.5b-v2 Minimal EN Corpus 2.5B 2.0 A deterministic English-language pretraining corpus for approximately 125M-parameter GPT-style models. This is the curated 2.0 generation of Minimal EN Corpus: source-pinned, model-free quality filtered, exact- and near-deduplicated, and decontaminated against the public evaluation splits of ARC, HellaSwag, PIQA, WinoGrande, OpenBookQA, BoolQ, MMLU, and GSM8K. The name refers to the original approximately 2.5B-token source mixture. After cleaning… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-2.5b-v2.text1M<n<10M0 likes326 downloads2d agoHugging FaceSlayerLab /gollem-corpus-16b-pl GoLLeM Corpus v4 PL — SlayerLab/gollem-corpus-16b-pl The exact pretraining corpus of the Polish base model GoLLeM v4 (250M, trained from scratch) — released before training, so the published bytes are byte-identical (sha-tied) to what the model will see. Successor of SlayerLab/gollem-corpus-2b-pl (the v2/v3 corpus), scaled ~7.5x with per-record provenance this time. 16.58B unique tokens (GoLLeM V32k tokenizer, measured) = ~1.96B curated + 14.62B cleaned Polish web. Per-record… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-16b-pl.texttext-generation1M<n<10M1 likes290 downloads11d agoHugging Face