CoolFace
Datasetpublic

placeholderlabs/pretrain-encyclopedic-mix

Normalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 7,321,112,013 (7.3B) Trainable tokens 7,321,112,013 (7.3B) Documents 6,498,683 Shards 58 UTF-8 bytes 27,544,569,761 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-encyclopedic-mix.

sourceHugging Faceupdated 15d agoView on Hugging Face
0likes89downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
placeholderlabs/pretrain-encyclopedic-mix · CoolFace