KristianS7/prepacked-fineweb-edu-llama2-32K-T2048
prepacked-fineweb-edu-llama2-32K-T2048 Pre-tokenized and BOS-aligned best-fit packed version of FineWeb-Edu for training with looped nanochat. Tokenized with the Llama 2 tokenizer (32,000 base vocab + 8 special tokens = 32,008). Stats Train split Source karpathy/fineweb-edu-100b-shuffle (1,821 shards) Total tokens 63.26B Total docs 97.1M Rows 30,873,598 Shards 2,059 (train-00000 to train-02058) Rows per shard ~15,000… See the full description on the dataset page: https://huggingface.co/datasets/KristianS7/prepacked-fineweb-edu-llama2-32K-T2048.
prepacked-fineweb-edu-llama2-32K-T2048
Pre-tokenized and BOS-aligned best-fit packed version of FineWeb-Edu for training with looped nanochat.
Tokenized with the Llama 2 tokenizer (32,000 base vocab + 8 special tokens = 32,008).
Stats
Train split
Val split
Common
Each row is a packed sequence of multiple BOS-delimited documents, cropped/filled to exactly 2,049 tokens. Rows are shuffled within each shard.
