placeholderlabs/pretrain-nemotron-math-mix-long-context
Normalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 1,446,296,439 (1.4B) Trainable tokens 1,446,296,439 (1.4B) Documents 42,379 Shards 23 UTF-8 bytes 4,965,563,314 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix-long-context.
068
