dlab-spp/corpus-verification
SPP Corpus Verification Checksums and document-boundary indices for verifying a rebuilt copy of the Synthetic Persona Pretraining (SPP) training corpus, byte for byte. The Megatron token streams themselves are 2.17 TB (annotated.bin 421 GB, compact.bin 1.75 TB) and are fully derived from the published reflections, the uid manifest, and the tokenizer recipe — so they are not published. These .idx sidecars carry per-document boundaries and lengths, which is enough to prove an… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/corpus-verification.
074
