glouriousgautam/lilm1-pretrain-mix-32b
LiLM Experiment 3 pretraining corpus Private research corpus with 32,000,010,072 globally exact-deduplicated train tokens plus 328,933,246 held-out tokens. Data are stored as EOS-delimited little-endian uint16 binaries with aligned Parquet provenance. This repository combines ODC-By FinePDFs-Edu, CC-BY-4.0 DCLM, ODC-By SmolLM/FineMath sources, Apache-2.0 UltraX subject to its upstream terms, per-file permissively licensed Stack-Edu code subject to The Stack v2 terms, StarCoder2… See the full description on the dataset page: https://huggingface.co/datasets/glouriousgautam/lilm1-pretrain-mix-32b.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face