CoolFace
Datasetpublic

placeholderlabs/pretrain-ultra-fineweb-mix

Normalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 73,328,911,347 (73.3B) Trainable tokens 73,328,911,347 (73.3B) Documents 92,923,076 Shards 578 UTF-8 bytes 364,562,557,837 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-ultra-fineweb-mix.

sourceHugging Faceupdated 16d agoView on Hugging Face
1likes1.1kdownloads
18 commits on main
85bff0116d ago

Quote the token count in short form beside the exact figure

sunnysetia
6db052616d ago

State the release size in tokens on the dataset card

sunnysetia
6e4224016d ago

Publish document-token mix 0fa7a48febfb

sunnysetia
1d0a4f416d ago

Add files using upload-large-folder tool

sunnysetia
efcd2ad16d ago

Add files using upload-large-folder tool

sunnysetia
de4485116d ago

Add files using upload-large-folder tool

sunnysetia
84857cb16d ago

Add files using upload-large-folder tool

sunnysetia
fa658b016d ago

Add files using upload-large-folder tool

sunnysetia
b57b30d16d ago

Add files using upload-large-folder tool

sunnysetia
869750716d ago

Add files using upload-large-folder tool

sunnysetia
233195416d ago

Add files using upload-large-folder tool

sunnysetia
08627a516d ago

Add files using upload-large-folder tool

sunnysetia
b3fdc9f16d ago

Add files using upload-large-folder tool

sunnysetia
a2d0a2516d ago

Add files using upload-large-folder tool

sunnysetia
1c15c1d16d ago

Add files using upload-large-folder tool

sunnysetia
43622ca16d ago

Add files using upload-large-folder tool

sunnysetia
cccf4e916d ago

Add files using upload-large-folder tool

sunnysetia
dc1b39516d ago

initial commit

sunnysetia