CoolFace
Datasetpublic

placeholderlabs/pretrain-nemotron-math-mix

Normalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 22,927,812,461 (22.9B) Trainable tokens 22,927,812,461 (22.9B) Documents 21,377,358 Shards 180 UTF-8 bytes 77,994,866,327 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix.

sourceHugging Faceupdated 15d agoView on Hugging Face
0likes154downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face