CoolFace
Datasetpublic

glouriousgautam/lilm1-pretrain-mix-32b

LiLM Experiment 3 pretraining corpus Private research corpus with 32,000,010,072 globally exact-deduplicated train tokens plus 328,933,246 held-out tokens. Data are stored as EOS-delimited little-endian uint16 binaries with aligned Parquet provenance. This repository combines ODC-By FinePDFs-Edu, CC-BY-4.0 DCLM, ODC-By SmolLM/FineMath sources, Apache-2.0 UltraX subject to its upstream terms, per-file permissively licensed Stack-Edu code subject to The Stack v2 terms, StarCoder2… See the full description on the dataset page: https://huggingface.co/datasets/glouriousgautam/lilm1-pretrain-mix-32b.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes773downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
glouriousgautam/lilm1-pretrain-mix-32b · CoolFace