CoolFace
Datasetpublic

glouriousgautam/lilm1-pretrain-mix-32b

LiLM Experiment 3 pretraining corpus Private research corpus with 32,000,010,072 globally exact-deduplicated train tokens plus 328,933,246 held-out tokens. Data are stored as EOS-delimited little-endian uint16 binaries with aligned Parquet provenance. This repository combines ODC-By FinePDFs-Edu, CC-BY-4.0 DCLM, ODC-By SmolLM/FineMath sources, Apache-2.0 UltraX subject to its upstream terms, per-file permissively licensed Stack-Edu code subject to The Stack v2 terms, StarCoder2… See the full description on the dataset page: https://huggingface.co/datasets/glouriousgautam/lilm1-pretrain-mix-32b.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes886downloads
Dataset Card

LiLM Experiment 3 pretraining corpus

Private research corpus with 32,000,010,072 globally exact-deduplicated train tokens plus 328,933,246 held-out tokens. Data are stored as EOS-delimited little-endian uint16 binaries with aligned Parquet provenance.

This repository combines ODC-By FinePDFs-Edu, CC-BY-4.0 DCLM, ODC-By SmolLM/FineMath sources, Apache-2.0 UltraX subject to its upstream terms, per-file permissively licensed Stack-Edu code subject to The Stack v2 terms, StarCoder2 documentation, Kaggle notebooks, StackOverflow, and GitHub issues subject to their respective upstream terms, and the canonical LiLM structured and tool corpora. Consult dataset_manifest.json for immutable revisions and per-shard hashes. Downstream users remain responsible for every upstream term and attribution requirement.

The corpus is globally exact-deduplicated over tokenizer IDs. It is not claimed to be globally near-deduplicated.