CoolFace
Datasetpublic

placeholderlabs/pretrain-web-mix

Normalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 73,646,210,143 (73.6B) Trainable tokens 73,646,210,143 (73.6B) Documents 61,059,647 Shards 590 UTF-8 bytes 341,537,872,441 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix.

sourceHugging Faceupdated 15d agoView on Hugging Face
0likes146downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
placeholderlabs/pretrain-web-mix · CoolFace