nikitastheo/v2-babylm-small-pol-ell-sequential_interleaved
0352
nikitastheo/v2-babylm-small-pol-ell-sequential_interleaved
Trained with train_clm.py, a Hugging Face Accelerate causal-LM training script (no Trainer).
Training details
- Base config:
model_configs/gpt_small_config.json - Tokenizer:
nikitastheo/babylm-pol-tokenizer - Max steps: 25340
- Learning rate: 0.0001
- LR scheduler: linear
- Warmup steps: 2534
- Batch size (per device): 32
- Gradient accumulation steps: 1
- Total train batch size: 32
- Language switch epoch: 10
