3B
Models
All models matching “3B”Datasets
All datasets matching “3B”copyright_gpt_neo_1_3Bforge-3b-pretrain-data
FORGE-3B Pretraining Data
Tokenized and packed pretraining data for the FORGE-3B language model.
Stats
Total tokens: 51.4070B
Domains: 10/10
Sequence length: 2048 tokens
Format: .npy shards of shape (N, 2048) with dtype uint32
Tokenizer: CRAYON (xerv-crayon, standard profile)
Domain Breakdown
Domain
Weight
Tokens (B)
Status
fineweb_edu
30%
15.0008
✓
thestack
16%
8.0011
✓
wikipedia
8%
4.2791
✓
openwebmath
8%
3.9654
✓
books
7%… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-pretrain-data.all_pile_gpt_neo_1_3Bstack-v2-starcoder2-3bllama-3b-residualsRULER-8192-Qwen2.5-3B-tokenizer
