UCLNLP/monoweb
0
MonoWeb Models
Pretrained language models released alongside the paper:
[The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining](https://arxiv.org/pdf/2601.00364)
Associated dataset: UCLNLP/monoweb-dataset
Model Details
All models are decoder-only transformers with 1.35B parameters, trained from scratch using the Llama-2 tokenizer (32K vocabulary). Architecture: 24 layers, hidden dimension 2048, 16 attention heads, context length 2048. Training was performed with Megatron-LM for ~143B tokens (34K steps).
Model Variants
Models are organized by language pair and training data configuration:
Each folder contains checkpoints saved every 2,000 steps from iter_2000 to iter_36000 (18 checkpoints per model).
