CoolFace
Datasetpublic

placeholderlabs/pretrain-nemotron-math-mix-long-context

Normalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 1,446,296,439 (1.4B) Trainable tokens 1,446,296,439 (1.4B) Documents 42,379 Shards 23 UTF-8 bytes 4,965,563,314 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix-long-context.

sourceHugging Faceupdated 15d agoView on Hugging Face
0likes68downloads
6 commits on main
1accc9c15d ago

Quote the token count in short form beside the exact figure

sunnysetia
e31f25415d ago

State the release size in tokens on the dataset card

sunnysetia
a12399916d ago

Publish document-token mix a40447ee0c3c

sunnysetia
a1ab7dc16d ago

Add files using upload-large-folder tool

sunnysetia
c1931f516d ago

Add files using upload-large-folder tool

sunnysetia
93ecbb816d ago

initial commit

sunnysetia