CoolFace
Datasetpublic

glouriousgautam/lilm1-pretrain-mix-32b

LiLM Experiment 3 pretraining corpus Private research corpus with 32,000,010,072 globally exact-deduplicated train tokens plus 328,933,246 held-out tokens. Data are stored as EOS-delimited little-endian uint16 binaries with aligned Parquet provenance. This repository combines ODC-By FinePDFs-Edu, CC-BY-4.0 DCLM, ODC-By SmolLM/FineMath sources, Apache-2.0 UltraX subject to its upstream terms, per-file permissively licensed Stack-Edu code subject to The Stack v2 terms, StarCoder2… See the full description on the dataset page: https://huggingface.co/datasets/glouriousgautam/lilm1-pretrain-mix-32b.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes773downloads

glouriousgautam/lilm1-pretrain-mix-32b · main · files are served by the source, never re-hosted here