CoolFace
Datasetpublic

placeholderlabs/pretrain-encyclopedic-mix

Normalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 7,321,112,013 (7.3B) Trainable tokens 7,321,112,013 (7.3B) Documents 6,498,683 Shards 58 UTF-8 bytes 27,544,569,761 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-encyclopedic-mix.

sourceHugging Faceupdated 15d agoView on Hugging Face
0likes71downloads
settings

This repository belongs to placeholderlabs on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namepretrain-encyclopedic-mix
visibilitypublic
licencenot set
gatedno
ownerplaceholderlabs
Account settings
placeholderlabs/pretrain-encyclopedic-mix · CoolFace