CoolFace
Datasetpublic

placeholderlabs/pretrain-encyclopedic-mix

Normalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 7,321,112,013 (7.3B) Trainable tokens 7,321,112,013 (7.3B) Documents 6,498,683 Shards 58 UTF-8 bytes 27,544,569,761 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-encyclopedic-mix.

sourceHugging Faceupdated 15d agoView on Hugging Face
0likes89downloads
8 commits on main
4168fdb15d ago

Quote the token count in short form beside the exact figure

sunnysetia
928354715d ago

State the release size in tokens on the dataset card

sunnysetia
c29770415d ago

Publish document-token mix aecff9e9f6ea

sunnysetia
ca6c2a215d ago

Add files using upload-large-folder tool

sunnysetia
4940ecd15d ago

Add files using upload-large-folder tool

sunnysetia
ae2cfe715d ago

Add files using upload-large-folder tool

sunnysetia
0e4abd815d ago

Add files using upload-large-folder tool

sunnysetia
53cfd0c15d ago

initial commit

sunnysetia