superbpe
Datasets
All datasets matching “superbpe”omega-5B-superbpe128k
Omega 5B SuperBPE 128k 70% FineWeb-Edu 15% StarCoder 10% FineMath 5% Gutenberg
Tokenizer: alisawuffles/superbpe-tokenizer-128k 128001 gigatoken GB/s
Tokens: 5047259136 seq_len 4096
finewebedu-superbpe-t180Kfineweb-edu10B-superbpe50304-3digit
FineWeb-Edu 10BT, tokenized with a two-phase 50,304-token SuperBPE vocabulary (3-digit number splitting)
All of HuggingFaceFW/fineweb-edu
sample-10BT, pre-tokenized in the binary shard format used by
modded-nanogpt, with a custom SuperBPE
vocabulary trained with BatchBPE (v2) on the
same corpus.
Phase 1 of the vocabulary split numbers into groups of up to 3 digits (\p{N}{1,3}). This is
the baseline in a planned comparison of first-phase digit-splitting rules, with sibling… See the full description on the dataset page: https://huggingface.co/datasets/alexandermorgan/fineweb-edu10B-superbpe50304-3digit.finewebedu-superbpe-t80Kfinewebedu-superbpefinewebedu-superbpe-t160K
