pre-tokenized
fineweb-edu-pretokenized-10K
Marin/Levanter Subsampled Pretokenized Dataset
Dataset
Train Urls:
gs://marin-us-central2/raw/fineweb-edu-c2beb4/3c452cb/huggingface.co/datasets/HuggingFaceFW/fineweb-edu/resolve/3c452cb
Factsheet
Original cache: gs://marin-us-central2/tokenized/fineweb-edu-24698d
Tokenizer: stanford-crfm/marin-tokenizer
Seed 42
Number of tokens: 10,390
(This readme is automatically generated by Marin.)
pretokenized-dolma
The Pretokenized Dolma Dataset
A pre-tokenized, pre-shuffled version of Dolma, the high-quality text corpus from AI2. This dataset is designed to be plug-and-play with the pico-train library.
Overview
Key Features:
Tokenized with allenai/OLMo-7B-0724-hf, a BPE-tokenized with a vocabulary size of 50280
Sequence length: 2049 tokens (2048 + 1 for next-token prediction)
Sharded into 10,000 Parquet files (~78MB each)
420B tokens total size (perfect for training a model for… See the full description on the dataset page: https://huggingface.co/datasets/pico-lm/pretokenized-dolma.rnacentral-pretokenized-shardedgemma-4-pretokenized-tracesdanish-pretokenized-16k
Danish Pretrain — pretokenized (byte-level BPE, 16k vocab) — v1 partial
Pretokenized version of jensjepsen/danish-pretrain
via jensjepsen/danish-tokenizer.
Note: This is a partial version — 94 of 96 shards. The 2 missing shards
crashed during tokenization (rare pathological content). See
jensjepsen/danish-pretokenized-16k-supp for the recovered rows.
93,200,464 docs
Schema: input_ids: list<int32>, attention_mask: list<int8>
94 zstd-compressed parquet shards
yeast-no-GTL-pretokenized-NT-v2
