CoolFace
Datasetpublic

joshycodes/op-spp-streams-v1

op-spp-streams-v1 — tokenized Megatron streams (SPP-format pretraining corpus) Tokenized Dolma v1.7 (ODC-BY) subsample in Megatron IndexedDataset format (uint16, SmolLM2 tokenizer + <assistant> extension from epfl-dlab/spp-training): compact dense-packed 2049-token windows; annotated/canary one document per window, EOD-padded. Built for a Synthetic-Persona-Pretraining-recipe run (arXiv:2608.13482) — these files carry ONLY document tokens (raw public-corpus text); the persona… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/op-spp-streams-v1.

sourceHugging Faceodc-byupdated 1mo agoView on Hugging Face
0likes35downloads
Dataset Card

op-spp-streams-v1 — tokenized Megatron streams (SPP-format pretraining corpus)

Tokenized Dolma v1.7 (ODC-BY) subsample in Megatron IndexedDataset format (uint16, SmolLM2 tokenizer + <assistant> extension from epfl-dlab/spp-training): compact dense-packed 2049-token windows; annotated/canary one document per window, EOD-padded. Built for a Synthetic-Persona-Pretraining-recipe run (arXiv:2608.13482) — these files carry ONLY document tokens (raw public-corpus text); the persona annotations ride in a separate private sidecar and are inserted at training time. See corpus_meta.json for the exact run arithmetic.