SlayerLab/minimal-en-corpus-2.5b-v2
Minimal EN Corpus 2.5B 2.0 A deterministic English-language pretraining corpus for approximately 125M-parameter GPT-style models. This is the curated 2.0 generation of Minimal EN Corpus: source-pinned, model-free quality filtered, exact- and near-deduplicated, and decontaminated against the public evaluation splits of ARC, HellaSwag, PIQA, WinoGrande, OpenBookQA, BoolQ, MMLU, and GSM8K. The name refers to the original approximately 2.5B-token source mixture. After cleaning… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-2.5b-v2.
This repository belongs to SlayerLab on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
